Conceptual

Hierarchical Token Compression for Long-Context Video Language Models

A hierarchical video-token compression approach that exploits the visual redundancy of long videos to compress the token sequence from clip level to video level, cutting the context length by roughly fifty times with almost no accuracy loss. This lets multimodal large language models process hours-long video (tens of thousands of frames) efficiently, and is paired with a short-to-long training curriculum and long-video benchmarks to build strong long-context video assistants.