Conceptual

Peer-Judged LLM Tournaments as Red-Teaming for Competitive Contract Negotiation

Introduces a round-robin evaluation protocol in which open-source generative language models negotiate the same sales contract head-to-head as buyer and seller, and the resulting agreement is scored by a panel of the remaining models acting as judges. Treating these competitive matchups as an adversarial red-teaming probe, the study measures which models systematically win favorable terms, exposes role-dependent biases (e.g., models that dominate as seller but fail as buyer) and safety/fairness vulnerabilities that isolated benchmarks miss, and derives practical model-selection guidance for LLM-mediated legal negotiation.