Most production agents have one validator. Usually a JSON schema check, sometimes a regular expression, and a line in the system prompt politely asking the model not to do the thing you are worried about. That is not a guardrail system. It is a single assertion with good intentions, and it fails predictably: it rejects malformed output and passes well-formed nonsense.
Validation is a layered system, and the layers are not interchangeable. Each can prove a different class of statement about an output, and almost all of the design work is deciding which statement each layer is allowed to settle.
The four layers, by what they can prove
Schema. Deterministic. Proves structure: fields, types, enum membership, numeric ranges, required combinations. Costs a couple of milliseconds. Says nothing about content.
Policy. Deterministic. Proves membership and permission against data you control: does this SKU exist, is this account inside the caller's tenant, is this URL on the allowlist, does this refund sit within the authority of the role that asked for it. Costs a few lookups, so tens of milliseconds. Its coverage is exactly equal to what you have enumerated. Not more.
Semantic. Probabilistic. Proves nothing; suspects usefully. Is every factual claim here supported by the retrieved context, is this the wrong register, is this a refusal dressed up as an answer. Costs a model call, so hundreds of milliseconds.
Human. Decides. Slow, expensive, and the only layer that can accept responsibility.
The ordering rule is the most useful sentence in this article: never let a probabilistic layer adjudicate a claim a deterministic layer can settle. I have reviewed more than one system in which an LLM judge was asked whether a response contained a valid product code. A lookup answers that exactly, in single-digit milliseconds, and is right every time. The judge is right most of the time, costs three hundred milliseconds, and gives you no way to find out which times it was not. Every question you can push down a layer, push down.
Layers therefore run in order of certainty, which usually coincides with order of cost and is not the reason.
Verdicts, not booleans
A validator that returns true or false is unusable in production, because you cannot route on it. Four outcomes, each carrying a reason code:
Pass.
Repair. The output is wrong in a way the system can fix without going back to the model: coerce a type, trim a field, drop an unsupported citation. Cheap, and it must be published as loudly as a rejection, because a climbing repair rate is a model regression wearing a disguise and a layer that silently tidies up is an effective way of not noticing one.
Retry. Send it back with the reason attached, and bound it. Two attempts, not "until it passes". An unbounded retry against a deterministic failure is an infinite loop with an invoice attached: the model cannot satisfy a constraint it was never able to satisfy, so it tries forever.
Escalate or degrade. Hand it to a person, or return the safe answer: the one that says less rather than the one that guesses.
The reason code is what makes everything after this possible. "Blocked" on a dashboard is noise. blocked: ungrounded_claim and blocked: entity_not_in_tenant are two different engineering problems with two different owners, and a layer that cannot tell you which one fired is not instrumented, however good its chart looks.
Placement: before the irreversible act, not at the end of the run
The common architecture puts validation at the end, on the final text. By then the agent has called three tools and two of them wrote something. Output validation on the last message is not a guardrail on an agent. It is a guardrail on a chat.
Every tool call with a side effect gets validated on its arguments, at the boundary, before the call, and that check is nearly always a policy check: does this entity exist, does the caller hold this authority, is this value inside the permitted band. Final-text validation still happens, and it is about grounding and register rather than safety, because by the time anyone reads the summary the money has moved.
(Whether a given step should be model-decided at all is a different and larger question. This article assumes the decision is the model's and asks how you check it.)
The semantic layer has one reliable job
The semantic layer is bad at "is this answer good". Ask a judge model whether a response is helpful and you get a number that moves with phrasing and position and bears no fixed relationship to anything a user would report.
It is considerably better at a narrower question: is every claim in this output present in the text that was supplied to the model? That is a comparison between two pieces of text you both hold. No world knowledge, no calibration against a hidden standard, and it degrades gracefully when unsure. Entailment against supplied context, register classification and refusal detection are the three jobs worth giving this layer. Anything requiring external truth belongs one layer down, with a lookup.
One thing to get right, and it is counter-intuitive enough to be worth testing rather than believing: use a different model family from the generator. A validator that shares the generator's lineage shares its blind spots, and it will agree with the generator precisely on the outputs where agreement is worthless. A cheap model from another family is usually a better grounding checker than a strong model from the same one.
Replay testing, and the corpus almost nobody builds
Here is the asymmetry that makes validation tractable in a way the rest of an agent stack is not. Validators are deterministic, so you can test them without the model.
Build a corpus of real bad outputs. Every time a person rejects something, a layer fires, or an incident produces a bad answer, store that output together with the input and the retrieved context that produced it, plus the policy version in force at the time. The context is not optional: a grounding validator cannot be replayed against an output alone, and a corpus of bare bad answers tests the schema layer and nothing above it.
That corpus is a unit-test suite. It runs in CI, in seconds, with no model calls and no non-determinism, so you can change a validator with the confidence of an ordinary code change. Nothing else in the pipeline offers that. It is not the golden set used for quality evaluation and it does a different job: the golden set asks whether the system got better, this asks whether a specific check still catches a specific thing.
Two measurements, and you need both:
- Catch rate against the bad corpus.
- False-positive rate against a held-out corpus of good outputs, sampled from real approved traffic.
A validator tuned on bad examples alone converges on rejecting everything, and it does so while its dashboard improves. This is the most common way a guardrail programme quietly damages a product: block rate rises, incident count falls, and nobody is measuring the good answers that stopped being delivered.
The corpus rots asymmetrically. Entries that a later fix made unreachable are still good tests; the point is that the validator would catch them again, so keep them. What goes stale is the labelling, and what stales it is a policy change, after which some old bad outputs are legitimate and the suite is quietly enforcing last quarter's rules. So every entry carries its label, who applied it and the policy version, and a policy change triggers a review of the affected slice rather than of everything. A few hundred well-labelled entries comfortably beat ten thousand unlabelled ones.
Fire rate as a first-class metric, and the two ways it lies
Every layer publishes its fire rate. A guardrail nobody has seen fire in a month is either unnecessary or broken, and the dashboard cannot tell you which.
Synthetic canaries resolve that. On a fixed schedule, inject a known-bad output into the validation path out of band and assert it was caught. Now a zero fire rate means the check is alive and the traffic was clean, which is information. Without canaries it means nothing, and a validator disabled by a configuration change sits behind a reassuring flat line for as long as you let it.

The second lie is the denominator. If your ordering means the semantic layer only sees the small share of outputs the cheap layers found suspicious, its fire rate is a fraction of a fraction and comparing it with the schema layer's is meaningless. Publish both: fires per layer invocation, which says whether the layer works, and fires per run, which says what users experience.
A rising fire rate is also not automatically a model regression. Segment by input class first. The commonest cause of a step change in block rate is a new customer asking a new kind of question, which is a coverage problem rather than a quality one and has a different fix.
The arithmetic, and what a timeout means
Layers in series add latency, and the semantic layer dominates. Take a generation at 2.5 s at the median, schema at 2 ms, three policy lookups at roughly 25 ms in total and a small-model grounding check at 400 ms. That is around seventeen per cent added at the median, and rather more at the tail, on every response. Measure the tail end to end rather than composing it from per-stage quantiles: stage tails do not line up in a way that lets you add them.
Two structural answers, and they are not equivalent. Stream into a buffer and release on validation, which costs perceived streaming and is the honest choice for anything a user will act on. Or run the expensive layer on the flagged subset plus a random sample of unflagged traffic. The sample is not optional: without it you have no estimate of what the cheap layers are missing, and the skipped path is unmeasured rather than safe.
Then decide, per layer and per action class, what a timeout means. Failing open under load fails at exactly the moment the system is least healthy. Failing closed turns a provider slowdown into an outage. Both are defensible; what is not is discovering your position inside a try block somebody wrote in a hurry. It belongs in configuration, under a name, and on the runbook.
What each layer cannot catch
Stated plainly, because a validation architecture sold as complete is worse than none:
- Schema cannot catch a well-formed lie, and most damaging outputs are well-formed.
- Policy cannot catch what you have not enumerated. Its coverage is a list; the list needs an owner and a review date, and every incident should end by asking whether it earns an entry.
- Semantic cannot give you a bound. A judge reporting 0.8 is not eighty per cent likely to be right, and thresholds do not transfer across prompt or model versions.
- Human cannot scale, and drifts. A reviewer at volume converges on the agent's answer.
None of the four is sufficient. The point of layering is that their failure modes are unlike each other, which is the only real defence anybody has.
What I got wrong
Two, both expensive.
I signed off a grounding validator that used the same provider and model family as the generator. The fire rate settled at about two per cent and looked healthy for weeks. It was healthy: the validator agreed with the generator's invented citations, confidently, because it found the same things plausible for the same reasons. The true rate surfaced only when a human sampling queue started reading approved outputs instead of blocked ones.
And I shipped a semantic layer that failed open on timeout, with alerting on fire rate and nothing on validator availability. A provider slowdown pushed the check past its deadline on most requests for several hours. Everything passed. The dashboard showed an unusually clean afternoon.
The short version
Four layers, ordered by what they can prove, and nothing probabilistic settles a question a lookup can answer. Verdicts with reason codes, and a bound on retries. Validate tool arguments before the side effect, not the summary afterwards. Give the semantic layer entailment against supplied context and run it on a different model family from the generator. Keep a corpus of real bad outputs with their inputs and policy version, and test against good outputs too. Publish fire rates per layer and per run, and prove the layers are alive with canaries. Decide what a timeout means before you find out.
Guardrails are not a model problem. They are ordinary engineering, testable in a way the model itself never is, which makes them the most reliable part of an agent stack rather than the vaguest.
Illustrations generated with AI.
