Conceptual

Interactive Benchmark for Language-Agent Experimental Design and Model Discovery

Evaluating an autonomous agent on the whole scientific cycle rather than one half of it: the agent must choose which experiment to run, observe the outcome, and revise its theory. Each environment is written as a generative probabilistic model, which makes two otherwise impossible measurements available - the expected information gain of a proposed experiment, computed with a nested Monte Carlo estimator, and a communication score obtained by handing the agent's natural-language explanation to a second agent that never saw the data and measuring how much its predictions improve. Students learn why coupling experimental design to model revision changes what an evaluation can detect, and why explanation-to-a-novice is a workable proxy for the quality of a discovered theory.