The first honest conversation about agent evaluation usually happens the week after a silent regression. Somebody changed a prompt, a provider shipped a model update, or a new document type entered the retrieval corpus, and quality moved in a direction nobody measured until a customer described it.
The reflex is to write tests. The reflex is right and the usual implementation is wrong, because the thing being tested does not behave like software. Run the same input twice and you get two different answers, both defensible. A suite that fails intermittently gets muted, as every team eventually does with a flaky test, and a muted suite is worse than none because it still appears in the pipeline.
Score the trace, not the answer
Most harnesses compare the final output to an expected one. For a single model call that is reasonable; for an agent it throws away the interesting half of the run. An agent can reach the right answer down a wrong path — wrong tool, error, retry, somewhere acceptable. Scored on output that is a pass. It is a case that got lucky, cost four times what it should have, and will fail the moment the recovery path changes.
The inverse is worse. A case fails, the aggregate drops two points, and an engineer spends a day reading model output when the cause was three steps upstream — a retrieval call that returned nothing because a filter was over-specified, after which no model could have answered.
So evaluate the whole trace — input, every tool call and response, final output, cost and wall-clock.
The hard part is what to assert against it, because most tasks have no single correct trajectory. Several paths are reasonable, and a set that treats one as canonical will fail your next real improvement. Assert on properties of the path rather than the path:
- Bounds. Step count under a ceiling, no repeated identical call, no more than one retry per tool. These survive replanning.
- Required and forbidden events. This class of task must touch the authorisation check; it must never call the write tool. Encode the contract, not the implementation — "the search tool was called" is worth asserting only where calling it genuinely is the contract, otherwise you have pinned today's plan into the set and made improvement look like regression.
- Outcome. What the user actually wanted.
A failure taxonomy pays for itself here: classify a failure as retrieval, tool-argument, planning or generation before a human opens it, and triage stops being archaeology. That is its own piece of work; the harness only has to store enough trace to make it possible later.
Build the golden set from production traces
A hand-written set encodes what the team imagined users would do. Traces encode what they did, which is consistently stranger.
Sample deliberately. A random thousand traces is mostly the easy majority case, so the aggregate ends up dominated by inputs that never break. Stratify by task type, so each capability can move its own number; by tool-call shape, because the three-tool and single-tool paths fail differently; by known failure class, oversampled against real frequency, because that is where regressions appear first; plus a proportional slice of the boring majority, so you can still tell if you have broken the common case.
Two or three hundred cases is enough to start and is a number a small team can curate. A retrieval golden set is the nearest relative you may already have, and it does not transfer.
Then version it rather than freezing it. Frozen and refreshed are each wrong alone: a set that never changes measures last year's product, and one that quietly gains cases produces scores you cannot compare across months. Use a frozen core that changes only by explicit decision plus a rotating margin of roughly a fifth, refreshed quarterly from recent traffic, retiring the oldest rather than accumulating. Every refresh bumps the version and re-baselines. A score means nothing unless you can name the set version, model snapshot and judge version behind it.
One warning: strip personal data at capture time, not at use time. An evaluation set is production data that lives in a repository, gets cloned to laptops and lands in CI logs.
Remove the variance you can, then measure what is left
Most of what makes an agent harness flaky is environmental and fixable. Record and replay tool responses — an external API returning different data on Tuesday tells you nothing about your agent. Pin the retrieval index, and break score ties deterministically on document ID. Pin the model to a dated snapshot rather than a floating alias, because undated family names are free to change weights, quantisation or routing underneath you. That matters more than temperature does.
What remains is irreducible, and temperature does not remove it. Greedy decoding removes deliberate sampling, not non-determinism: identical requests at temperature zero can return different tokens because the provider batches your request with other traffic and floating-point addition is not associative, so the same computation in a different batch tie-breaks differently. Mixture-of-experts routing that depends on batch composition adds more of the same. You do not control the batch, so you do not control the output.
So run each case more than once and make the per-case pass rate the unit. Four things to be honest about:
- Five repeats gives the order of magnitude of your noise, not a precise standard deviation: estimated from five values, the error on that estimate is around a third of it. Five tells you whether the floor is one point or six. If a threshold carries real weight, measure it once at twenty or thirty repeats.
- You gate on a difference between two runs, and the spread of a difference of two independent runs is about 1.4 times the spread of one. A threshold read off single-run spread is tighter than it looks.
- Do not import anyone else's thresholds. Aggregate noise depends on your set size and repeat count: a number right for 250 cases at five repeats is wrong by a factor of two for sixty.
- Declare one gate number in advance. Check the aggregate, every tier and every criterion against its own threshold nightly and something crosses a line most nights by chance — the muted suite again. One gate; the rest are review signals.
This also answers the cost objection: once fixtures and a pinned index remove the environmental variance, repeats only pay for model and judge variance. Thirty-case smoke set per pull request, full set nightly, repeats only at the release gate.
Use the cheapest oracle the property allows
If a property can be checked by code, check it by code. Does the JSON validate? Does the total match? Were the tool arguments drawn from the allowed enum? Cheap, and they never drift — provided they run against fixtures rather than a live system. Asserting a customer ID exists by querying the production database gives you a test that breaks when somebody deletes a record.
Push as much into that category as the task allows, and revisit the boundary, because what feels subjective is often checkable once stated precisely. "Is the tone appropriate" is not assertable; "contains no hedge phrase from this list" is, and catches most of what that question was asking.
Reserve the model-as-judge for what is left — completeness, whether an explanation answers the question actually asked — and treat it as a non-deterministic component inside your harness. Three rules make one usable.
Pin the judge model and version, and never let it float. A judge that updates is indistinguishable, in the aggregate, from a system that regresses — the most confusing failure a new harness produces.
Keep a human-labelled anchor set and measure agreement. When the judge model or rubric changes, score the anchor set against the human labels before believing anything the judge says afterwards. Use a chance-corrected statistic — Cohen's kappa or equivalent — because raw agreement on an imbalanced label flatters badly: ninety per cent agreement on a ninety-per-cent-one-class label is what a coin that always says yes gives you. At fifty cases only large drops are detectable, so size the set for the comparison you want.
Score named criteria as binary judgements, not a one-to-ten quality score. Judges cluster their scale and the same rubric produces different distributions across versions, so aggregate judge scores travel badly. Binary per-criterion judgements are portable and tell you what moved. The cost is real: a case near the decision boundary flips zero-to-one rather than drifting two-tenths, so per-case variance rises as comparability improves. For regression detection that trade is the right way round.
Evaluate every tier against the whole set
If you route between a cheap model and an expensive one, the risk is not mainly dilution in the aggregate. It is that for work the cheap tier does not currently serve, you have no measurement of that tier at all.
That is what makes a routing change dangerous. A cost-saving push moves a class of work down to a tier never scored on it, and the failure appears in production on a day nobody deployed a model change — which makes it very hard to attribute.
So run the full set against each tier independently, including cases it does not serve. You are not asking whether the system is good; you are asking what each tier is safe for, and that table is your routing policy — and the day a cheaper model launches, the answer to what it can take over is a day's work rather than a fortnight's.
What to gate a deploy on
Not the aggregate — that tells a human to look. Gate on three things, and be honest that each carries a judgement made in advance rather than none.

The regression list. Cases that were passing and are not now. Under per-case pass rates that needs a written rule: five-of-five dropping to three-of-five is a regression, five to four usually is not. One named case an engineer can open beats any mean.
Safety assertions, where a single failure blocks. The tempting phrase here is "no statistics", and it is wrong — one run is one sample, and passing once does not establish that a stochastic path is safe. Run the safety subset at higher repetition and treat a failure in any repeat as a block.
Cost and latency budgets. Cost per task is stable enough to gate directly, and an agent that retries its way to a correct answer at forty times the token cost has not succeeded. Latency is less obliging, since CI runs against a shared provider under variable load: gate it on a statistic you have a floor for, or on step and token counts, which are yours.
Everything else belongs in the report a human reads before approving.
What decides whether any of this lasts
Golden sets rot in two directions. They go stale, still measuring the product you shipped last year. And they get overfit, because engineers debug against the cases in the set, so those improve while everything outside them does not and the score rises while the system does not.
Core-and-margin versioning handles the first. The second needs a holdout: a slice nobody debugs against, scored only at release. The moment someone opens a holdout case to fix it, it moves into the main set.
What I got wrong
I have shipped a harness that scored only the final answer, and spent the better part of a week diagnosing a quality drop that turned out to be a retrieval filter change. The model was fine throughout, and the trace that would have shown it in ten minutes was not being stored.
And I have compared an aggregate score across a judge model upgrade and read the movement as a system regression. We spent two days looking for a change we had not made. Pinning the judge version and keeping an anchor set is now the first thing built, before a single case is scored, because everything downstream is uninterpretable without it.
The short version
Assert on properties of the path, not on a canonical path. Mine the set from real traffic and version it as a frozen core plus a rotating margin. Fixture the tools and pin the index and a dated model snapshot before you measure anything, then measure what noise is left. Use the cheapest oracle each property allows. Score every tier against the whole set, and gate on named regressions rather than on a mean.
The harness is a fortnight of engineering. The labelled set behind it takes real work, and with that discipline it is one of the few artefacts in an agent system worth more next year than today.
Illustrations generated with AI.
