Conceptual

Video-Centric Multimodal Textbook Corpus for VLM Pretraining

A high-quality image-text interleaved pretraining corpus built from 22,000 class-hours of instructional videos rather than web crawl. Keyframes, ASR transcripts, and OCR text are extracted and arranged in temporal order to give dense foundational knowledge, coherent context, and tight image-text alignment; VLMs pretrained on it gain knowledge/reasoning performance (ScienceQA, MathVista) and stronger few-shot interleaved in-context ability.