2501.00601
DreamDrive generates photo-realistic, 3D-consistent driving videos along an ego vehicle's trajectory for scalable self-driving training, unifying the two prior families that each fall short: reconstr…
Synthesizing 3D-consistent, photo-realistic driving videos along an ego vehicle's trajectory by combining generative and reconstruction methods. A video diffusion model produces generalizable visual references from in-the-wild street-view data, which are elevated to a 4D (spatial-temporal) scene via a hybrid Gaussian representation: static (time-independent) Gaussians for the background plus dynamic (time-varying) Gaussians for moving objects, decomposed self-supervised without object annotations. Rendering the 4D scene with Gaussian splatting yields controllable novel-view driving videos with temporal and multi-view consistency. Exemplified by DreamDrive, evaluated on nuScenes and in-the-wild data for perception and planning.
DreamDrive generates photo-realistic, 3D-consistent driving videos along an ego vehicle's trajectory for scalable self-driving training, unifying the two prior families that each fall short: reconstr…