Simulated-Annealing Fine-Tuning for Multilingual Counterspeech Generation
A method for generating hate-speech counterspeech in low-resource languages that applies a simulated-annealing optimization schedule while fine-tuning a multilingual language model, steering it toward respectful, specific, and factually grounded responses. It demonstrates that a stochastic global-optimization schedule can improve generation quality where labeled data is scarce, and it examines the reliability of LLM-as-judge evaluation for the counterspeech task.
2501.00713
Counterspeech is text written to counter and neutralize hate speech while preserving free expression, but authoring it by hand does not scale. This paper presents CODEOFCONDUCT, a context-aware syste…