J
Jazz
Text
A controlled, component-wise study of automatic LLM benchers (frameworks that rank language models by alignment with human preference). It decomposes a bencher into four choices - input set, evaluation model (LLM-as-judge), evaluation type (pointwise vs pairwise), and aggregation method (e.g. Elo) - and gives evidence-based recommendations for each. Key findings: bencher reliability drops sharply when ranking similarly-performing models, and an evaluation model's instance-level judgment accuracy does not predict its system-level ranking effectiveness.