Conceptual

Spatial-Temporal Object Referring in Video Language Models

Extending Video Large Language Models from holistic video understanding to object-level (referring) understanding: a user designates any object by a pixel-level mask/click at a timestamp, and the model reasons about that object's attributes, actions, and relationships throughout the video. Covers the spatial-temporal object encoder design (per-object regional tokens via mask pooling plus temporal aggregation of an object's tokens across frames), object-level video instruction data, and benchmarks that separately score spatial description and temporal/relational reasoning. Exemplified by the VideoRefer Suite (VideoRefer-700K dataset, VideoRefer model, VideoRefer-Bench).