Generating a Chinese Counterspeech Corpus with an LLM-as-Judge and Simulated Annealing
The pipeline behind PANDA, the first Modern Standard Mandarin counterspeech dataset: filter hate speech with an LLM discriminator, generate candidate counterspeech from an ensemble of LLMs, search the candidate space with simulated annealing whose Boltzmann acceptance uses an LLM-as-a-judge quality score, rank finalists by round-robin tournament, then human-verify — plus the documented limits of LLM-as-judge scoring in a non-Eurocentric language.
2501.00697
Chinese-language counterspeech resources — constructive replies that rebut hate speech without censorship — were essentially nonexistent. This paper introduces PANDA, the first Modern Standard Mandar…