Conceptual

Generative 4D Driving Scene Modeling with Hybrid Gaussians

Synthesizing 3D-consistent, photo-realistic driving videos along an ego vehicle's trajectory by combining generative and reconstruction methods. A video diffusion model produces generalizable visual references from in-the-wild street-view data, which are elevated to a 4D (spatial-temporal) scene via a hybrid Gaussian representation: static (time-independent) Gaussians for the background plus dynamic (time-varying) Gaussians for moving objects, decomposed self-supervised without object annotations. Rendering the 4D scene with Gaussian splatting yields controllable novel-view driving videos with temporal and multi-view consistency. Exemplified by DreamDrive, evaluated on nuScenes and in-the-wild data for perception and planning.