Conceptual

Human Evaluation and Rank Correlation in NLP

Human evaluation collects ratings from people to judge machine-generated text on qualities like surprise or quality, and rank-correlation statistics such as Spearman's rho measure how well an automatic metric agrees with those human rankings. It is the standard way to validate that a proposed metric tracks human judgment.