The first agentic feature I put in front of real users cost about what we modelled. The second one cost four times the forecast, and the forecast was not wrong about price per token or about traffic. It was wrong about a thing that does not exist in a chat product: an agent decides how much work to do.

That is the structural difference, and most cost surprises in agentic systems trace back to it. In a chat product, a request is a model call. Cost per request is roughly a constant, so cost is a traffic problem, and product teams already know how to forecast traffic. In an agentic system, a request is a budget the agent spends at its own discretion — a planning call, some tool calls, an observation loop, a retry when a tool returns something malformed, a summarisation pass when the context grows past the window. The number of steps is data-dependent. A small minority of your traffic will be doing a disproportionate share of the reasoning, and you will not know which requests until you look.

So the useful mental model is not price per token. It is: what is the most this task is allowed to spend, and what happens when it tries to spend more?

What the bill actually looks like

Take a document-analysis agent — the sort of thing that sits behind "summarise this file and tell me what changed." Model it the naive way and you get: one request, one frontier-model call, a few thousand tokens in, a few hundred out. Fine.

Ship it and instrument it, and the trace for a single request looks closer to this:

StepCallsNotes
Intent classification1Cheap, small model, easy to forget in the forecast
Planning1Frontier model, full system prompt
Tool calls (retrieval, parse, compare)3–6Each one re-sends accumulated context
Retry / repair0–2Malformed tool output, schema violation, timeout
Context compaction0–1Fires once the window fills
Final synthesis1Frontier model, largest input of the run

Six to twelve model calls where the cost model said one. But the call count is the smaller half of the problem. The larger half is that in most agent loops the whole working conversation is re-sent as input on every step, so cumulative input tokens grow with the square of the step count rather than linearly. A twelve-step run does not cost twelve times a one-step run. As a worked example rather than a measurement: if each step appends a similar amount and none of the prefix is cached, twelve steps sum to roughly seventy times one step's input, not twelve. The real multiple depends on how much each step appends and how much of the prefix your provider will let you cache. That is where the factor-of-four forecast miss came from, and it is why "we'll optimise the prompt" is not the fix — trimming the system prompt by 200 tokens saves those tokens once per step, and the step count is what is hurting you.

Stepped chart of the input sent at each of twelve agent steps when every step re-sends the conversation so far and nothing is cached. A band of one step’s input per step adds up to 12 times; the re-sent conversation rises above it step by step, for roughly 70 times one step’s input in all. A worked example from the text. Every step re-sends the run so far input tokens sent, per step Roughly 70× one step’s input over twelve steps, not 12× step 1 step 12 One step’s input at each step: 12× in all Steps above it: the run so far, re-sent Worked example from the text, not measured: similar appends, no prefix cached.
The step count is what hurts, not the length of the prompt. Worked example from the text, not a measurement.

Two corollaries worth stating plainly, because both cost me money before I believed them:

Failures are not free — they are your most expensive traffic. A run that burns nine calls and then fails validation has spent full price for zero output. In the early Continuum360 work, failing runs were a low single-digit share of volume and a meaningfully higher share of spend, because failures correlate with hard inputs, and hard inputs are exactly the ones that loop. Any cost dashboard that only shows cost per successful task is hiding this.

Two balance scales side by side. One hangs level, a small heap of pale cubes against a single finished block; the other is tipped hard by a much taller heap of orange cubes against an empty pan, showing a failed run that spent the most and produced nothing.
A failed run pays full price for zero output. A dashboard that only shows cost per successful task hides it.

Your p99 is a cost metric, not just a latency metric. The tail of your step-count distribution is the tail of your bill. If you monitor mean tokens per task you will watch a flat line while a small population of pathological runs eats the margin.

Layer one: token budgets as a first-class runtime object

The single highest-leverage change is to stop treating the token budget as a number in a planning spreadsheet and make it an object that travels with the request and gets decremented.

Concretely, a task arrives with a ceiling denominated in money rather than raw tokens, and every model call in the run debits it at that call's own rate. Money is the right unit because tokens are not fungible: output tokens are priced well above input tokens with every major provider, the tiers differ from each other by roughly an order of magnitude, and a cached input read is cheaper again than an uncached one. A flat token count stops meaning much the moment you route across tiers. The orchestrator reads the remaining balance before each step and changes behaviour as it depletes. That last part is what distinguishes a budget from a limit: a limit only tells you when to stop, whereas a budget lets you degrade deliberately.

Degradation ladder that has held up for me, in order of what to give up first:

  1. Drop optional enrichment. The nice-to-have retrieval pass, the second-opinion check, the formatting polish. Cheapest to lose, least visible to the user.
  2. Reduce breadth before depth. Retrieve five chunks instead of twelve. Compare two documents instead of six. Users notice a shallower answer far less than a wrong one.
  3. Downgrade the tier. Finish the run on the mid-tier model rather than the frontier one. Quality drops; the task completes.
  4. Compact aggressively. Summarise the working context hard, accepting some fidelity loss, to stop the quadratic growth. Budget for the fact that rewriting the context invalidates any prompt cache built on it, so the step after a compaction is dearer than the ones around it.
  5. Stop and hand off. Return partial output with an honest statement of what was not done, and offer escalation.

Note what is absent: silently truncating the input and answering anyway. That converts a cost problem into a correctness problem, and a confidently wrong answer costs more than the run you refused to finish. If the budget cannot cover the task, say so.

Diagram of a task budget, held in money, draining as the run proceeds, beside the five things the run gives up in order: optional enrichment, breadth before depth, the model tier, context through aggressive compaction, and finally it stops and hands off with partial output. Silently truncating the input is not on the ladder. What to give up as the budget runs down Balance in money, read before every step 1 Drop optional enrichment cheapest to lose, least visible to the user 2 Reduce breadth before depth five chunks instead of twelve 3 Downgrade the tier finish on the mid tier; quality drops 4 Compact aggressively stops the growth; invalidates the cache 5 Stop and hand off partial output, and say what was not done Not on the ladder: silent truncation a cost problem becomes a correctness problem Schematic. Set the ceiling per task type.
A limit only tells you when to stop. A budget, read before every step, lets the run give things up in order and say so when it cannot finish.

Set ceilings per task type, not globally. "Summarise this paragraph" and "reconcile these two documents" have no business sharing a budget.

Layer two: model-tier routing that survives contact with reality

Routing is the other large lever, and the reason it disappoints people is usually that they attempt it as one decision at the front door.

The front-door version — classify the incoming request, send it to a cheap, mid or frontier model — works, but it is graded on the accuracy of a classifier that is looking at the request before any work has been done. It cannot know that this particular document is malformed, or that retrieval will return contradictory passages. Route conservatively and you pay frontier prices for easy work; route aggressively and you get quality complaints that cost more to fix than the savings.

What works better is routing per step, with escalation, and treating the tiers as roles rather than as a quality ranking:

  • Small model, structural work. Classification, extraction against a schema, routing decisions, reformatting, deciding whether a tool result is well-formed. This is a surprisingly large share of the calls in a typical agent loop and almost none of the value-add reasoning. Moving it off the frontier model is the cheapest win available and rarely changes output quality at all.
  • Mid tier, the default execution lane. Most steps in most runs. It should be the presumption, not the fallback.
  • Frontier tier, earned by need. Planning on genuinely ambiguous inputs, final synthesis where the reasoning is the product, and any step the mid tier has already failed.

Escalation needs an explicit trigger, and the honest ones are behavioural rather than predictive: the mid-tier output failed schema validation; a self-consistency check across two samples disagreed; the step has already been retried once; a confidence signal came back below threshold. All of these are observable after cheap work, which is exactly why they beat a front-door guess.

Then the important interlock: escalation must be budget-aware. An escalation path with no ceiling is a mechanism that converts your hardest traffic into your most expensive traffic automatically, and hard traffic is not evenly distributed across customers. Cap escalations per task. When the cap is hit, fail visibly.

Diagram of three model tiers as roles: the small model does structural work such as classification and extraction, the mid tier is the default lane for most steps, and the frontier tier is earned by need. A step moves up from mid to frontier only on observed evidence, such as a failed schema check, and escalations are capped per task. Tiers as roles, with capped escalation Frontier: earned by need planning on ambiguous inputs, final synthesis Escalate on evidence, after cheap work: schema check failed, samples disagree, already retried once, low confidence Capped per task: at the cap, fail visibly Mid tier: the default lane most steps in most runs Small model: structural work classify, extract, route, check tool output Routed per step, not once at the front door.
Triggers you can observe after cheap work beat a guess at the front door. An escalation path with no ceiling turns your hardest traffic into your most expensive.

Two limitations I would not paper over. First, tier routing spreads your quality surface across several models, so you now need per-tier evaluation — a regression that only shows up on the small model is easy to miss for weeks. Second, mixed-tier output has a consistency cost: the same question can get differently-shaped answers depending on which lane it took, and for some products, that variance is a worse problem than the bill. If your product promises consistency, pay for one tier and save elsewhere.

Layer three: not paying twice for the same thought

Caching in agentic systems is worth more than in chat products, because agent loops are repetitive by construction.

Three distinct mechanisms, often conflated:

Prompt caching, where the provider supports it. Structure the prompt so that the stable part — system instructions, tool definitions, retrieved context that will not change during the run — sits at the front, unchanged, and the variable part goes last. The mechanism is a prefix match, so the cache helps only as far as the first token that differs from the previous call — one edit near the front costs you everything behind it — and entries expire on a short idle timer. Nor is a hit free: it is discounted input, not zero, and on some providers writing the entry in the first place costs a premium over an ordinary input token. It is still the largest single lever on the re-sending problem, and the first thing to check when a cost model comes in high, because a prompt assembled in the wrong order defeats it entirely.

Two prompts side by side. With the stable part first, the system instructions, tool definitions and retrieved context are read from the cache and only the variable part is paid in full. With one edit near the front, the cache stops at the edit and everything behind it is paid in full. Cached up to the first token that differs Stable part first One edit near the front System instructions Tool definitions Retrieved context Variable part System instructions Tool definitions Retrieved context Variable part Hit on the stable prefix variable part at full price Miss from the edit onwards all behind it at full price A hit is discounted input, not free.
Put the stable part first and the variable part last. A prompt assembled in the wrong order defeats the cache entirely.

Exact-match caching on deterministic sub-steps. Tool results, parsed documents, retrieved chunks for identical queries. Boring, reliable, no semantic risk.

Semantic caching on near-duplicate requests. Real savings in high-volume, narrow-domain surfaces where users genuinely ask the same thing in different words. Also the one that will hurt you, and its failure mode is worth being blunt about: set the similarity threshold too loose and you return the answer to a neighbouring question, with full confidence and no signal to the user or to you that a substitution happened. That is a near miss rather than a stale entry, it does not show up in your error rate, and in an assessment or compliance context it stops being a cost story at all. My rule is that semantic caching is for advisory, informational surfaces, never for anything with a decision or a record attached.

Alongside caching, the structural saving is simply not putting a model where a function belongs. In every agent codebase I have reviewed, including my own, there are steps calling a language model to do something deterministic — validating a date, sorting a list, checking a numeric range. Each one is cheap. Collectively they are a line item, and they are also a latency and reliability line item, so removing them pays three times.

What to instrument, if you instrument nothing else

Cost control in agentic systems is an observability problem before it is an optimisation problem. The four numbers I would not run without:

  1. Cost per completed task, by task type — not per request, and not averaged across types.
  2. Step-count distribution, at p50 and p99 — the gap between them is your exposure.
  3. Spend on failed runs, as a share of total spend — the number nobody instruments and everybody is surprised by.
  4. Tier mix per task type — because tier mix drifts. A prompt change that makes the mid-tier model marginally less reliable will quietly shift traffic to the frontier lane, and the bill moves before anything else does.

If a cost dashboard cannot answer "which task type got more expensive this week, and was it price, volume, step count or tier mix?", it will tell you that costs went up and nothing more.

The trade-off worth naming

None of this is free. Budgets, routing, escalation triggers and per-tier evaluation are real orchestration complexity, and the complexity lands in the hardest part of the system to debug. On a low-volume internal tool, that trade is bad — pay the frontier-model bill and spend the engineering time elsewhere.

The reason to build it is that agentic cost has a different shape from the costs product teams are used to. It scales with reasoning, and reasoning scales with the difficulty of the input, which means your unit economics are set by your hardest users rather than your average ones. You cannot forecast your way out of that. You can only bound it.

Bound it early, while the traces are simple enough to read.

Illustrations generated with AI.