There is a debugging ritual I have watched play out on almost every RAG system I have been close to. A user reports a wrong answer. The team reads the answer, winces, and blames the model. Someone swaps in a bigger model, or rewrites the prompt to say "only use the provided context" a third time, slightly louder. The answer changes shape and stays wrong.

When we finally instrumented this properly on one of our production pipelines, the pattern was blunt: in well over half of the reported failures, the evidence needed to answer correctly was never in the context window at all. The model was not hallucinating in any interesting sense. It was doing what a language model does when asked to answer from documents it was never shown: producing something fluent and plausible. The failure happened two stages earlier, in retrieval, and nobody was looking there because nobody could see it.

Generation failures are loud; a bad answer is right in front of you. Retrieval failures are silent, because the model papers over the missing evidence with confident prose. If you only evaluate the end of the pipeline, you will spend months tuning the one component that was mostly working.

Diagram of a RAG pipeline in five stages: ingestion, chunking, embedding, retrieval and generation. Retrieval is marked as the silent failure, where the evidence never reaches the context. Generation is the loud failure, a bad answer you can see. Where RAG answers go wrong Ingestion Chunking Embedding Retrieval Generation Silent failure the evidence never reaches the context Loud failure a bad answer you can see
Teams debug where the failure is visible, not where it happened.

Build the harness before you tune anything

The fix is unglamorous: a retrieval evaluation harness, built before you touch chunk sizes, embedding models or rerankers. Without one, every change is a vibe. With one, every change is a measurement.

The harness needs three things.

A golden set of real queries. Mine them from production logs, not from your imagination. Synthetic queries written by the team who built the corpus have a fatal property: they are phrased the way the documents are phrased, which makes retrieval look far better than it is. Real users ask short, ambiguous, misspelt questions with the key entity missing. Pull a few hundred of them, covering both the head (frequent, boring) and the tail (rare, revealing). On our side, roughly 400 queries has been enough to make regressions visible without making maintenance a project of its own. Be honest about what a set that size can resolve: the sampling noise on a recall figure sits at a couple of points, so a one-point move is not a result.

Relevance labels. For each query, a domain expert marks which chunks in the corpus actually contain the answer. This is the expensive part and there is no honest way around it: budget a few minutes per query of someone senior's time, and treat the labelled set as an asset with an owner, a version number and a review cadence. Teams that skip labelling end up using LLM-as-judge for relevance, which is workable for triage but drifts with the judge model and quietly inherits its blind spots. Use it to pre-filter, not as ground truth.

A metric you actually act on. For RAG, the primary number is recall@k, where k is the number of chunks you actually pass to the model. Score it per query as a hit or a miss — did the evidence needed to answer land in those k chunks — rather than as the fraction of all relevant chunks retrieved, which is the textbook definition and much harder to act on. The question the hit rate answers is the only one that matters at this stage: did the evidence make it into the window at all? If recall@5 is 0.6, two out of five user questions were unanswerable before the model saw a single token, and no amount of prompt engineering will recover them. Mean reciprocal rank is worth tracking as a secondary signal because position within the window still influences what the model attends to, though it only ever sees the first relevant chunk and so tells you nothing about the second piece of evidence a multi-part question needs. Precision matters mainly as a cost metric: irrelevant chunks are tokens you pay for, and past a point they crowd out the evidence you did retrieve.

What the harness changes

Once the harness exists, arguments that used to run for a week resolve in an afternoon.

Chunking stops being folklore. We had the standard debate: small chunks for precision versus large chunks for context. Run against the golden set, the answer on our corpus was neither of the defaults. Moving from uniform fixed-size chunks to structure-aware chunking that respected section boundaries moved recall@5 from 0.61 to 0.78, a larger gain than any embedding model swap we ever measured. Your corpus will differ. That is exactly the point: the harness tells you about your corpus, not about the blog post you read.

Hybrid retrieval stops being a checkbox. Dense embeddings are genuinely weak on identifier-heavy queries: product codes, statute numbers, error strings, people's names. These sit in the tail of the query distribution, which is why demos never surface them and production always does. A BM25 leg alongside the dense index, fused with something as simple as reciprocal rank fusion, exists precisely for that slice. The harness will show you the size of the slice before you commit to running two indexes forever.

Rerankers get bought with evidence. A cross-encoder reranker earns its latency only in one situation: recall@50 is high while recall@5 is low, meaning the evidence is being found but drowned. Then reranking the top 50 into a better top 5 is worth the added latency, and that latency is a number to measure on your own hardware and candidate count rather than to take from anyone's blog post. If recall@50 is already poor, the reranker has nothing to rescue and you have bought latency for nothing. I have seen the reranker added first in both situations; the harness is how you know which one you are in.

Two schematic charts of recall against k. In the first, recall at k = 5 is low but recall at k = 50 is high, and the gap between them is highlighted: the evidence is found but drowned, so a reranker can close the gap. In the second, recall is poor even at k = 50, so there is nothing for a reranker to rescue. Does a reranker have anything to rescue? recall@k k = 5 k = 50 recall@k k = 5 k = 50 Found, but drowned recall@50 high, recall@5 low: a reranker can close the gap Never found recall@50 already poor: nothing to rescue Schematic; k is on a log scale.
A reranker earns its latency only when the evidence is already in the top 50.

Version everything, then wire it into CI

An embedding model is not a stateless dependency, and this is where scale stops being a footnote. Change it and every stored vector is stale; re-embedding a large corpus is a real bill and hours of pipeline time, so the decision deserves a measurement in front of it. We learnt this the way everyone does. A minor embedding-model upgrade, announced as an improvement, dropped recall on our tail queries by eight points while leaving head queries untouched. The average barely moved; averages are where regressions hide. In Continuum360, the orchestration layer we build at Longshort Labs, that eval run is now a hard gate: no retrieval config change, no embedding upgrade, no chunking revision ships without a harness pass against the versioned golden set.

A glass tank of perfectly still water whose level surface hides a cluster of sharp orange rocks in one corner of an otherwise smooth floor, showing how an average can stay flat while one group of queries gets worse.
A level average can sit on top of a regression in the tail.

Treat the trio as one versioned unit: index build, embedding model, golden set. An eval score means nothing unless you can say which of the three it was measured against.

Monitoring in production, where there are no labels

You cannot compute recall live, because live traffic has no relevance labels. What you can watch are proxies that move when retrieval degrades: the distribution of retrieval similarity scores, which is only comparable with the embedding model held fixed (a downward drift suggests the corpus and the queries are growing apart), the share of queries whose best match falls below your similarity floor, and the rate of answers where the model declines or hedges. None of these proves failure on its own; together they tell you when to pull a fresh sample of production queries, label them, and fold them into the golden set. That loop, sampled labelling feeding the eval set, is the only thing that keeps the harness honest as usage shifts. A golden set nobody refreshes becomes a museum of last year's questions.

Diagram of a loop. Live traffic has no relevance labels, so there is no live recall. When proxies move (similarity drift, the share of queries below the similarity floor, hedging answers), a fresh sample of production queries is labelled by a domain expert and becomes the next version of the golden set, versioned with the index and the model. The loop then repeats. Keeping the golden set honest Live traffic no relevance labels, so no live recall Proxies move similarity drift, below-floor share, hedging Sample and label fresh production queries, by a domain expert Golden set, next version versioned with the index and the model
A golden set nobody refreshes measures last year’s questions.

What this does not solve

Honesty about the limits, because the failure modes move once retrieval is fixed. High recall does not guarantee a grounded answer; a model can ignore or distort evidence it was correctly given, and that needs its own groundedness evaluation downstream. The harness also cannot rescue a bad corpus: if the answer exists in no document, retrieval evaluation will faithfully report that you are perfectly retrieving nothing. And the labelling cost is real and recurring, not a one-off. I know of no serious deployment that has escaped it.

But the ordering stands. Retrieval evaluation first, generation tuning second. It is the difference between debugging the component that is actually failing and shouting at the one that merely delivers the bad news.

Illustrations generated with AI.