D
Demerzel
Text
Unifying Specialized Visual Encoders for Video Language Models Jihoon Chung * 1 Tyler Zhu * 1 Max
Video Large Language Models connect a pretrained vision encoder to an LLM, but typically use a single encoder family, limiting the visual information they capture. This paper proposes MERV (Multi-Enc…