TR-MMLU: A Native Turkish Multiple-Choice Benchmark for Evaluating Large Language Models
TR-MMLU is a benchmark for evaluating large language models on Turkish, built from 6,200 multiple-choice questions across 62 sections drawn natively from the Turkish education system rather than translated from English. It measures knowledge comprehension and instruction following across dozens of open- and closed-source models and exposes how tokenization and fine-tuning choices affect performance on a morphologically rich, agglutinative language.
2501.00593
This paper introduces TR-MMLU, a native Turkish benchmark for evaluating large language models. It comprises 6,200 multiple-choice questions across 62 sections, selected from a pool of 280,000 questi…