Conceptual

Vision-Only 3D Scene Understanding for Vision-Language Models

A visual prompting paradigm that gives 2D Vision-Language Models 3D spatial understanding of indoor scenes from video alone: a Bird's Eye View image reconstructed from the video supplies global layout, and Spatial-Temporal Object markers assign consistent object IDs across frames and the BEV to establish global-local correspondence, the ingredient VLMs lack for 3D reasoning.