Conceptual
Login

Generating a Test Dataset for a Prompt Evaluation

Assemble the inputs an eval runs against — drawn from real traffic, written by hand, or model-generated — covering the edge cases and failure shapes the prompt is meant to survive.