Conceptual
Login

Evaluating LLM Agents in Murder-Mystery Social Deduction Environments

A benchmark design in which large language model agents play a scripted murder-mystery game to measure social reasoning under incomplete and adversarial information. Agents are assigned Culprit or Civilian roles and move through an Open Conversation phase, an Interaction phase of Ask and Investigate actions, and a Murder Voting phase, supported by auxiliary summarization, suspicion-tracking, trust-modelling and rerun modules. Four complementary metrics score the transcripts: Trust Inclination Index, Clue Investigation Capability, Interactivity Capability Index and Script Compliance Index, with a separate neutral judge model used for scoring after an ablation showed a judge shares a bias with models from its own family. The concept teaches how to turn an open-ended social game into a reproducible evaluation harness and how to read the resulting scores as evidence about reasoning, deception handling and instruction adherence rather than as a single leaderboard number.