S
Sonny
Text
Generating a Test Dataset for a Prompt Evaluation — Building with the Claude API
Assemble the inputs an eval runs against — drawn from real traffic, written by hand, or model-generated — covering the edge cases and failure shapes the prompt is meant to survive. Taught in the cour…