Conceptual

Calibrated Multidimensional LLM Evaluation with Personalized Rubrics

A method that turns a large language model into a calibrated, multidimensional text evaluator. Human-written rubric questions are answered by the LLM as probability distributions over options, and a small feed-forward network with judge-independent and per-judge parameters is trained on human annotations to predict each rater's scores, including an overall-quality judgment, cutting the error of the raw LLM judge roughly in half.