Component-Balanced Measurement of Gender Stereotype Bias in Language Models
Shows that intrinsic gender-stereotype bias benchmarks such as StereoSet and CrowS-Pairs each emphasize only some facets of stereotyping, so their bias scores correlate weakly, and that re-curating and rebalancing benchmark data across social-psychology stereotype dimensions substantially improves agreement between measurement approaches. Students learn how a benchmark's data distribution drives its apparent bias score, and how grounding evaluation data in a structured framework of stereotype components yields more consistent intrinsic bias measurement.
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets Mahdi
This paper argues that intrinsic gender-stereotype bias benchmarks for language models (StereoSet and CrowS-Pairs) each capture only partial facets of gender stereotyping, so their bias scores correl…