The invoice arrives before the architecture does
The pattern is consistent enough to set a watch by. A team ships an LLM feature, usage is modest, the bill is a rounding error. Six months later the feature is popular, the bill is the third-largest line in the cloud spend, and someone senior asks for a plan. The plan that comes back is almost always "switch to a cheaper model" — which is the last lever you should reach for and the one most likely to cost you the product.
The good news is that the savings available are large and mostly structural. Caching, batching, routing and quantisation together cut managed API spend by 50–90% on typical production workloads without touching model quality. The bad news is that most teams apply them in the wrong order, and the wrong order is expensive.
This is where the money goes, and the order we work in.
First: know what you are actually paying for
Before any optimisation, instrument. You cannot route traffic you cannot classify.
The unit that matters is cost per resolved task, not cost per token or even cost per request. A request that costs four cents and answers the question is cheaper than three requests at one cent each that end in a handoff to support.
The dimensions worth carrying on every span:
gen_ai.request.model which model
gen_ai.usage.input_tokens prompt size
gen_ai.usage.output_tokens generation size
gen_ai.usage.cached_tokens what the cache absorbed
app.task_type classification, extraction, chat, agent…
app.tenant who to attribute it to
app.resolved did this end the task
Two findings fall out of this on nearly every codebase, and both are worth more than any model swap:
Input dominates. Typical products run 10:1 or worse input to output. The prompt, the retrieved context, the tool definitions, the conversation history — that is the bill. Teams instinctively optimise the response length, which is the small half.
A few task types own the spend. Rarely more than three. Once you can see spend by task_type, the plan writes itself, because you stop optimising the eleven things that together cost nothing.
Second: cache, because the cheapest call is the one you skip
Two distinct mechanisms, routinely confused, and you want both.
Prompt caching
Provider-side. The stable prefix of your prompt — system instructions, tool definitions, a long document — is cached and charged at a steep discount on subsequent requests. It reduces API cost by 45–80% and improves time-to-first-token by 13–31%.
Getting it to hit requires one discipline: stable prefixes. Cache keys are prefix matches, so anything volatile near the front of the prompt destroys the hit rate for everything after it.
# Destroys the cache: the timestamp changes the prefix on every call.
system = f"You are a support agent. Current time: {now()}. {TOOLS}"
# Hits: volatile content moves after the stable block.
system = f"You are a support agent. {TOOLS}"
messages = [{"role": "user", "content": f"[time: {now()}]\n{user_text}"}]
Audit for this specifically. We have found timestamps, request IDs, A/B flags and randomised few-shot examples sitting in prompt prefixes — each one silently paying full price for a block that should have cost a tenth.
Semantic caching
Yours. Store request–response pairs and return them for semantically similar queries, eliminating the inference call entirely. On FAQ-shaped distributions this reaches up to 68% call reduction.
It is also the one that will hurt you if you are careless, because a false hit returns a confidently wrong answer to a question nobody asked. The guardrails we insist on:
- Scope the key. Embed the query, but partition by tenant, locale and any authorisation context. A cross-tenant cache hit is a data breach, not a cost saving.
- Set the threshold high, then lower it with evidence. Start at 0.95 cosine. "Reset my password" and "reset my payment method" are closer in embedding space than anyone's intuition suggests.
- Never cache personalised or time-sensitive answers. If the response contains account state, it is not cacheable. Mark this at the task-type level rather than trusting per-call judgement.
- Log hits as hits. A cache hit that produced a bad answer must be traceable, or you will spend a week blaming the model.
Semantic caching goes first because applying it before routing means you build the routing system for the traffic that survives, which is a much smaller and more interesting problem.
Third: route what is left
Now, and only now, model selection. The framing that works is not "which model is best" but "what is the cheapest model that passes the evals for this task type."
That sentence has a dependency: you need the eval harness. Routing without an eval harness is guessing, and it fails in the worst way — quality degrades slowly, nobody notices for a quarter, and the fix is a rollback nobody can justify.
A three-tier structure covers most products:
| Tier | Handles | Typical share |
|---|---|---|
| Small | Classification, routing, extraction, short factual answers | 60–70% |
| Mid | Multi-step reasoning, summarisation, most chat | 20–30% |
| Large | Code generation, long-context synthesis, hard reasoning | 5–15% |
The routing decision itself must be cheap relative to what it saves. The latency budget is real: rule-based routing adds under 1ms, embedding-based about 5ms, and heavier ML classifiers 50–100ms, against typical LLM responses of 500–2,000ms.
Start with rules. Task type, input length and a small set of keyword signals get you most of the benefit at zero measurable latency. Graduate to a learned router only when you can show rules are misrouting enough to matter.
Route up on failure, not just down on confidence. The pattern that keeps quality intact:
result = small.run(prompt)
if result.confidence < 0.7 or result.refused or validator.rejects(result):
result = mid.run(prompt) # escalate, don't fail
The escalation rate is a metric you watch. Rising escalation means your small tier is being asked to do something it cannot, and the fix is usually the routing rule rather than the model.
Fourth: the boring structural levers
Cheaper per token than anything above, and consistently neglected.
Trim the context. Full conversation history on every turn is the most common unforced error in the category. Summarise past a threshold and carry the summary. On long chat sessions this alone routinely halves input tokens.
Trim retrieval. Teams retrieve twelve chunks because twelve felt safe. Measure it: on most corpora, answer quality plateaus around five and the other seven are pure cost. This is an eval question with a definite answer.
Trim tool definitions. Sixty tools in every system prompt is thirty thousand tokens per turn. Scope the catalogue per agent and per turn.
Batch what is not interactive. Offline classification, enrichment, backfills — provider batch endpoints run around half price for work nobody is watching. The only requirement is being honest about which of your workloads are actually interactive. Usually fewer than claimed.
Making spend a property of the system
Optimisation decays unless it is enforced. The architectural move is to put a gateway in the path — the same gateway that handles tool access, if you have one — and give it four responsibilities:
app ──▶ gateway ──▶ provider
├─ semantic cache (skip the call)
├─ router (pick the tier)
├─ budget enforcement (per tenant, per task)
└─ attribution (cost per task, per tenant)
Budget enforcement is the piece that turns cost from a monthly surprise into a runtime property. Per-tenant ceilings, per-task-type ceilings, and a circuit breaker that degrades to the small tier rather than serving errors when a limit is approached. A runaway agent loop is a possibility in every agentic system; the question is whether it costs you forty dollars or forty thousand before anyone notices.
Attribution matters for a less obvious reason: it is what makes usage-based pricing possible. If you cannot say what a customer costs to serve, you cannot price them, and you will discover your worst-margin accounts by accident.
The gateway does not need to be exotic. Several open-source options do this off the shelf — Portkey went fully Apache-2.0 in March 2026 with conditional routing, circuit breakers and semantic caching built in — and building your own is a two-week project you should only take on if you have a specific reason.
The order, and why it is that order
- Instrument. Cost per resolved task, by task type and tenant.
- Cache. Prompt caching for the prefix, semantic caching for the repeats.
- Route. Cheapest model that passes the evals, escalating on failure.
- Trim. Context, retrieval depth, tool definitions.
- Batch. Everything nobody is waiting on.
- Only then, consider self-hosting or quantisation.
That last point deserves its own warning. Self-hosting an open-weights model is the optimisation teams reach for first because it feels like the real engineering. It is a genuine win at sustained high volume, and a trap below it — you are taking on GPU capacity planning, inference server tuning, model updates and an on-call rotation to save money you could have saved with a cache. Run the arithmetic against a fully-loaded engineer-month before anyone provisions a cluster.
The bar
You have a cost architecture, rather than a cost problem, when four things are true: you can state the cost of your top three task types without opening a spreadsheet; a runaway loop is bounded by a budget rather than by someone noticing; a model deprecation is a routing-table change; and finance can attribute spend to customers without a data project.
The cheapest token remains the one you never send. Almost everything above is an elaboration of that.