Conceptual

Multi-Encoder Video Representations for Video Language Models

Improving Video Large Language Models by fusing features from several heterogeneous visual and video encoders (e.g. DINOv2, ViViT) rather than a single backbone family. The features are spatio-temporally aligned, projected into a unified structure, and fused with cross-attention, giving the VideoLLM a richer, complementary video representation with minimal extra parameters and parallelizable, faster training.