Conceptual

Automatic Evaluation and Benchmarking of Large Language Models

Methodologies for scoring LLM outputs, including multiple-choice benchmarks, reference-based metrics, and LLM-as-judge approaches, and the biases (fluency, position, likelihood) that make evaluating open-ended reasoning hard.