An agent says it created the report. The final message is clear, the formatting is good, and the screenshot looks convincing.
The report does not exist.
This is a useful evaluation case because a language-only judge might still reward the answer. The system performed the appearance of completion without establishing the fact of completion.
For AI agent evaluation, I want the success condition to live outside the model’s own account of what happened. A report has an artifact identity and a persisted revision. A scheduled task has a durable record. A sent message has an acknowledged or explicitly uncertain delivery outcome.
The final sentence is evidence of what the model said. It is not proof of what the system did.
Evaluate the whole task, then isolate the cause
An agent application contains several systems that fail differently: a model, a context builder, a tool layer, a policy engine, a scheduler, external services, and a user interface.
A single end-to-end score can tell you that performance changed. It often cannot tell you why.
I would separate deterministic contract tests, integration scenarios, and model-dependent task evaluations. Contract tests establish invariants such as “a stale owner cannot commit a result.” Integration scenarios establish behavior across stateful components. Model evaluations test whether the system chooses and completes appropriate work on representative tasks.
Anthropic’s discussion of agent evals distinguishes the task, trial, grading process, transcript, and outcome.[1] That separation is useful because it stops the transcript from becoming the only thing you inspect.
Turn invariants into executable assertions
Start with claims that should not depend on model quality.
A cross-tenant artifact reference must be rejected. An expired approval must not dispatch. A cancellation from an old run must not cancel a newer one. A failed compaction must not advance the history boundary. A duplicate completion must not publish a second result.
These assertions can often be tested with ordinary unit and integration tests. You do not need a model call to find out whether a conditional database update checks the generation number.
Keep the failure fixtures small enough to understand. A test with fifty mocked services can become its own source of false confidence. For critical concurrency behavior, include tests against the actual persistence primitive you rely on, not only an in-memory imitation.
A mock that implements a lock more strongly than your real datastore is a particularly expensive form of optimism.
Design adversarial timelines, not only adversarial prompts
Prompt injection matters, but many serious failures require no malicious text.
A worker loses its lease while an external request is in flight. A human approves a document and someone edits it before dispatch. A child task finishes after its parent is cancelled. A browser command arrives after control was handed to a person. A stream reconnects after missing an update.
Write those timelines down and make the test harness control the relevant boundaries. Pause before a commit. Inject a timeout after the destination records an effect. Deliver completion events twice. Advance a fake clock past the approval expiry.
For each scenario, define both the backend invariant and the visible user outcome. “No duplicate action” is not enough if the UI falsely says the action definitely failed. Correct uncertainty is part of the product’s correctness.
Use task-specific evidence for completion
A useful task rubric contains an observable target.
For document analysis, that may be a set of supported claims with citations to authorized evidence. For a chart, it may be the correct dataset, units, aggregation, and published revision. For a browser workflow, it may be a verified page state or an external receipt, not a generated description of the expected page.
The evaluation harness should retrieve this evidence independently of the final answer. It can then compare the answer against the actual outcome.
Partial success deserves a defined representation. A task that produced a correct draft but correctly stopped before an unauthorized send is not equivalent to a task that never found the source document. Whether it passes depends on the requested workflow and the policy, not on whether the model sounded apologetic.
A judge is an instrument with its own failure modes
Model-based grading can help with open-ended qualities such as relevance, clarity, and whether an explanation addresses the question. It should not decide every invariant.
Use deterministic checks for exact values, access decisions, artifact existence, schema validity, and required evidence where possible. Use bounded rubrics for judgment-heavy dimensions. Calibrate the grader against human-reviewed examples, including plausible but wrong answers.
Hide irrelevant provider labels from the judge where feasible. Vary answer order in comparisons. Include examples where verbose, confident prose should lose to a shorter, correct result.
When the judge and a deterministic check disagree, do not average away the problem. Investigate which property each measured. A beautiful explanation of a nonexistent artifact should not receive a passing aggregate score because the writing was strong.
Reliability is a distribution, not one successful replay
A task that passes once may fail on another trial because of model sampling, scheduling, or external timing.
Run repeated trials where that variability matters and report the conditions. Keep the number of attempts, model configuration, tool versions, and task set visible. A pass rate without its sample size and trial definition is hard to interpret.
Separate first-attempt success from success after repair. Both can be useful, but they have different latency and cost implications. Also track failures that are safe and recoverable separately from failures that violate an authority or data-isolation invariant.
A rare safety violation should not disappear inside an average quality score. Some conditions are release gates, not weighted preferences.
Optimize cost per verified completion
Tokens per request can fall while the system becomes less useful. A shorter context may omit evidence, cause retries, or send the agent down a longer tool path.
A more decision-useful measure is the total cost of the evaluated workload divided by verified successful tasks, alongside the failure breakdown. Include model calls, relevant execution costs, repairs, and abandoned attempts within a clearly stated accounting scope.
Do not hide expensive failures by calculating cost only for the successful subset. Also do not let a single aggregate conceal different workload classes. A simple document lookup and a multi-step report should have separate views before you combine them.
Track latency in the same spirit. Time to the first token is not time to a usable result. For agent workflows, the latter is often the number the user actually cares about.
Keep the dataset useful without leaking private work
Build synthetic fixtures that preserve the failure shape rather than copying customer documents or production traces into an evaluation repository.
A stale-approval test needs a resource that changes between review and dispatch. It does not need a real customer’s contract. A context test needs conflicting instructions and evidence boundaries. It does not need a confidential conversation.
Separate development examples from held-out evaluation tasks. Record fixture provenance and revisions. When a failure becomes a regression case, avoid tuning only to its wording; add variations that preserve the underlying difficulty.
Do not automatically turn every user interaction into training or evaluation material. Access, consent, retention, and the intended data use still apply.
Compare models through supported workflows
A provider swap should trigger a capability-oriented regression suite, not just a syntax check on the API adapter.
Test structured output, tool-call sequencing, context projection, refusal behavior, repair, cancellation handling, and the application’s interpretation of usage data. A model that excels at analysis may still require a different workflow for reliable artifact editing.
Use explicit support levels: supported, supported with a constrained path, or unsupported. A fallback that removes an approval requirement to keep the task running is not graceful degradation.
Pin the configuration used in each evaluation record. “The model got worse” is not a useful diagnosis when the prompt, tool catalog, compaction rule, and renderer changed at the same time.
A release report should explain the decision
The final artifact of evaluation should be a short, inspectable release record: what changed, what was tested, where performance improved, which failures remain, and why the release is acceptable.
That record is more valuable than a dashboard full of green averages nobody can explain.
The goal is not to prove the agent will never fail. It is to make the important failure modes concrete, keep unacceptable ones from being normalized, and know whether the next change actually made the system better.
A demo shows what the agent can do. An evaluation system shows what you are prepared to claim it can do reliably.
Technical notes
[1] Anthropic, Demystifying evals for AI agents. The test scenarios and release framework in this article are proposed application practices, not reported benchmark results.
Continue reading
Deterministic Infrastructure for AI Agents: Make Every Action Accountable · Agent Context Compaction: Treat Summaries as Checkpoints