Conceptual

Feed-Forward Transformer Reconstruction of Dynamic 3D Gaussian Scenes

Reconstructing dynamic (moving) 3D scenes from sparse posed images with a single feed-forward Transformer pass rather than slow per-scene optimization. The model predicts per-frame 3D Gaussians plus a per-Gaussian velocity, and aggregates Gaussians across timesteps via self-supervised scene flow to transport them to a target time, giving complete ('amodal') novel-view renders at any moment. Trained only with reconstruction losses, it separates static and dynamic content and yields emergent motion masks without motion supervision. Exemplified by STORM for large-scale outdoor/driving scenes.