Hierarchical Token Compression for Long-Context Video Language Models
A hierarchical video-token compression approach that exploits the visual redundancy of long videos to compress the token sequence from clip level to video level, cutting the context length by roughly fifty times with almost no accuracy loss. This lets multimodal large language models process hours-long video (tens of thousands of frames) efficiently, and is paired with a short-to-long training curriculum and long-video benchmarks to build strong long-context video assistants.
2501.00574
VideoChat-Flash: methods for long-context video understanding in multimodal large language models. Its core contribution is Hierarchical video token Compression (HiCo), which exploits the heavy visua…