The Production Agent Evaluation Playbook
A practical scorecard for task success, tool accuracy, recovery behavior, cost, and human escalation.
Evaluate the job, not the conversation
A fluent answer can still fail the business task. Start by describing the completed job in observable terms: the right record was found, the correct tool was called, the policy was followed, and the user reached a valid outcome.
Separate task success from presentation quality. This prevents a polished response from hiding a broken workflow.
Build a scenario set from reality
Use real support tickets, operations cases, and failure reports. Include ordinary requests, ambiguous inputs, missing data, tool outages, permission boundaries, and adversarial behavior.
- Happy paths with known outcomes
- Boundary cases and incomplete context
- Tool timeouts and malformed responses
- Requests that require human approval
- Prompt injection and unauthorized data requests
Score the system in layers
A single aggregate score is hard to diagnose. Track retrieval quality, decision quality, tool execution, response quality, and end-to-end completion separately. When the headline score changes, the layer metrics explain why.
Measure recovery, not just first attempts
Production agents will meet unavailable APIs, stale records, and unclear requests. Test whether the system retries safely, asks for missing information, or escalates with useful context instead of looping.
Set budgets before launch
Define acceptable latency, model spend, tool calls, and escalation rates per workflow. Budgets turn optimization into an engineering decision and expose regressions before they reach the invoice.
Operate a weekly evaluation loop
- Add real failures to the scenario set
- Review false passes and false failures
- Compare candidate prompts and models on the same suite
- Require regression checks before releases
- Share a short scorecard with product and operations owners