2501.00584
Addresses online (streaming, real-time) video understanding with multimodal large language models, where a model must perceive, memorize and reason over a continuous video stream without access to fu…
An approach to online (streaming, real-time) video understanding with a multimodal large language model, in which the model must perceive, remember and reason over a continuous video stream without seeing future frames. Its novel contribution is the Pyramid Memory Bank, which retains key spatiotemporal information from the stream at multiple granularities within a bounded memory budget, combined with an offline-to-online learning paradigm (an interleaved dialogue format plus a tailored instruction-tuning dataset) that adapts an offline video MLLM to the streaming setting. The resulting VideoChat-Online model is more efficient yet outperforms prior offline and online models, and is evaluated with the accompanying OVBench benchmark of past/current/future temporal tasks.
Addresses online (streaming, real-time) video understanding with multimodal large language models, where a model must perceive, memorize and reason over a continuous video stream without access to fu…