Conceptual

Omni Superpoint Transformer for Point-Cloud 3D Large Multimodal Models

A unified 3D vision-language model (3D-LLaVA) that consumes only point clouds and connects them to a large language model through the Omni Superpoint Transformer, a single module that selects visual tokens, encodes interactive visual prompts, and decodes text-conditioned 3D segmentation masks. Enables 3D question answering, captioning, and referring segmentation without offline preprocessing or task-specific heads.