Building an Evaluation Set
Explore how to build evaluation sets that test Claude systems against specific criteria such as task accuracy, safety, output format, and latency. Understand the importance of real, edge, and adversarial cases, and how to use automated grading methods to ensure system reliability and catch regressions before deployment.
A Claude system that looks good in a demo has passed one test, chosen by the person giving the demo. Real inputs are messier, and every change to a prompt, a model, or a tool can quietly break behavior that used to work. Without a way to measure the system, the team learns about these breaks from customers.
An eval set is a collection of test cases with known correct outcomes, run against the system to measure how well it performs. It turns a claim such as the agent is reliable into a number that can go up or down. In a system design interview, the question how would you know it works is almost guaranteed, and an eval plan is the expected answer.
Success criteria come first
An eval measures the system against success criteria, so the criteria have to exist before the cases do. Good criteria are specific and measurable. A goal such as the assistant should answer well cannot be tested, while a goal such as the assistant should pick the correct category for at least 95 percent of tickets can.
Most systems need criteria on several dimensions at once.
Task accuracy: How often the system reaches the correct answer or decision.
Safety behavior: Whether it refuses, escalates, or ignores what it should, such as instructions hidden in a document. ...