Conceptual

LLM-as-Judge Grading Calibrated Against Human Labels

grading with a model is how evals scale, and an uncalibrated judge is a confident random number generator