2501.00321
OCRBench v2 is a large-scale bilingual (English and Chinese) benchmark for evaluating the Optical Character Recognition (OCR) abilities of Large Multimodal Models (LMMs). It assesses eight core text-…
A benchmark methodology for evaluating the optical-character-recognition abilities of large multimodal models across many text-centric tasks and real-world scenarios at once. It spans eight core competencies — text recognition, referring, spotting, relation extraction, element parsing, mathematical calculation, visual text understanding, and knowledge reasoning — using human-verified question-answer pairs and metrics that expose recurring model weaknesses in fine-grained perception, layout understanding, complex-element parsing, and logical reasoning.
OCRBench v2 is a large-scale bilingual (English and Chinese) benchmark for evaluating the Optical Character Recognition (OCR) abilities of Large Multimodal Models (LMMs). It assesses eight core text-…