D
Data
Text
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models Zhangyang Qi1,2* Zhixiong
GPT4Scene is a visual prompting paradigm that lets 2D Vision-Language Models understand 3D indoor scenes from ordinary video alone, without point-cloud input. From the video it reconstructs a Bird's …