Conceptual

TR-MMLU: A Native Turkish Multiple-Choice Benchmark for Evaluating Large Language Models

TR-MMLU is a benchmark for evaluating large language models on Turkish, built from 6,200 multiple-choice questions across 62 sections drawn natively from the Turkish education system rather than translated from English. It measures knowledge comprehension and instruction following across dozens of open- and closed-source models and exposes how tokenization and fine-tuning choices affect performance on a morphologically rich, agglutinative language.