The integration argument is over
Two years ago every team building agents wrote its own tool layer. Function schemas in one shape for OpenAI, another for Anthropic, a third for whatever internal router had been stood up, and an adapter file that nobody wanted to own. That argument is settled. The Model Context Protocol is the wire format, the registry counts close to ten thousand servers, and 41% of surveyed software organisations report MCP servers in limited or broad production.
Winning the format argument did not make the problem easy. It moved it. What used to be "how do I describe this tool to the model" is now "how do I let an agent acting for a specific human touch forty internal systems without becoming the widest-reaching piece of infrastructure in the company."
That is the article. Not what MCP is — the spec is short and you should read it — but what falls over when you put it in front of real users, and the architecture we ship to stop it.
Three failure modes, in the order you will meet them
Every MCP rollout we have worked on has hit the same three walls, roughly in this order.
1. Auth propagation
This is the big one, and it is the top reported integration blocker from enterprise pilots. The naive deployment gives the MCP server a service account. The server holds a token for Jira, another for the data warehouse, another for the ticketing system, and it uses them for every request regardless of who asked.
You have just built a confused deputy. Every user of the agent now has the union of every permission the service account holds. A support engineer asks the agent to "check the account status for this customer" and the agent — helpfully, correctly, within spec — runs a query the support engineer could never have run themselves.
Nothing in the protocol stops this. It is an architecture decision, and the default is wrong.
2. Tool-definition bloat
Each MCP server advertises its tools. Attach eight servers and you are shipping perhaps sixty tool definitions, each with a name, a description and a JSON Schema, into the system prompt of every single turn.
We have seen this reach 30,000 tokens before anyone noticed, because it does not fail. It just quietly makes every request slower, more expensive, and worse: past roughly forty tools, selection accuracy falls off. The model starts picking search_documents when it wanted search_tickets, because the descriptions were written by different teams and both say "search for relevant items."
3. No blast radius control
An agent with a delete_record tool will eventually call delete_record. Not because the model is malicious, but because a user said something ambiguous and the trajectory went somewhere nobody tested. If the only thing between the model and your production database is a well-worded tool description, you do not have a control — you have a hope.
The gateway is not optional
The single highest-leverage decision is to stop letting agents talk to MCP servers directly. Put a gateway in the path, and make it the only thing that holds credentials.
user ──▶ agent runtime ──▶ MCP gateway ──▶ MCP servers ──▶ systems
(no creds) (policy, (per-request (Jira, DW,
identity, identity) tickets…)
audit, budget)
Concretely the gateway owns five jobs:
- Identity exchange. It converts the caller's identity into a downstream credential scoped to that caller, per request.
- Tool filtering. It decides which subset of the catalogue a given agent, in a given context, is even allowed to see.
- Policy. Allow, deny, or require human approval, evaluated per tool call rather than per session.
- Audit. One log line per tool invocation with the human, the agent, the tool, the arguments and the result hash.
- Budget. Token and call ceilings, enforced before the call rather than discovered on the invoice.
None of this is exotic. It is an API gateway with the vocabulary changed. The mistake teams make is treating MCP as a client-side concern because the first tutorial ran the server on a laptop.
Solving auth properly
The pattern that works is token exchange — RFC 8693 — with the gateway as the exchange point. The agent runtime never holds a downstream credential. It holds a token that says who the user is, and the gateway trades that for a narrower one.
# Gateway: one exchange per tool call, scoped to the caller and the tool.
async def downstream_token(user_token: str, tool: ToolRef) -> str:
return await idp.exchange(
subject_token=user_token, # who is actually asking
subject_token_type=ACCESS_TOKEN,
audience=tool.resource, # e.g. "warehouse.internal"
scope=tool.required_scopes, # least privilege, per tool
requested_token_type=ACCESS_TOKEN,
)
Three properties matter here:
- The token is audience-bound. A credential minted for the warehouse cannot be replayed against the ticketing system, so a compromised MCP server leaks one integration rather than all of them.
- Scopes are per tool, not per server. An MCP server exposing both
read_ordersandrefund_ordershould not hand the same token to both. - Expiry is short. Minutes. The exchange is cheap and cacheable within a turn; there is no reason to mint hour-long tokens for an operation that takes 400ms.
If your identity provider does not support token exchange, the fallback is a gateway-held mapping from user identity to a per-user stored credential. It is worse — you are now a credential store, with everything that implies — but it is still enormously better than one service account.
The one thing you should not do is pass the user's original token straight through to the MCP server. It is over-scoped by construction, and you have just handed a full-privilege credential to a process whose whole job is executing model-chosen instructions.
Keeping the tool catalogue small
The fix for tool bloat is not a bigger context window. It is admitting that tool selection is a retrieval problem.
We ship a two-stage pattern:
Stage one — static scoping. The gateway knows which agent is calling and which tools that agent is permitted. A support agent sees ticketing and account-read tools. It never sees deployment tools, so it cannot pick them, and they cost nothing in context.
Stage two — dynamic selection. Within the permitted set, if there are still more than roughly thirty tools, embed the tool descriptions once at startup and retrieve the top-k against the current turn.
# Only the tools this agent may use, narrowed to what this turn is about.
allowed = policy.tools_for(agent_id, user.roles) # 60 → 22
selected = tool_index.search(turn_text, within=allowed, k=12)
Two things make this work in practice, and both are unglamorous:
- Write the descriptions as a set, not individually. If two tools could plausibly match the same sentence, the descriptions are wrong. We keep a single owned file of tool descriptions and review them the way you would review an API's public documentation, because that is what they are — except the consumer is a model that cannot ask a clarifying question.
- Always include the tools the agent used in the last two turns. Retrieval that drops a tool mid-task produces a spectacular class of failure where the agent forgets it can do the thing it was just doing.
Below about fifteen tools, skip stage two entirely. Retrieval adds latency and a failure mode; a short, well-written catalogue beats a clever selector.
Blast radius: classify every tool before it ships
Give every tool a tier at registration time. This is a five-minute exercise that prevents the incident.
| Tier | Example | Control |
|---|---|---|
| Read | search_tickets, get_order | Log only |
| Write, reversible | add_comment, create_draft | Log, per-session rate limit |
| Write, irreversible | refund_order, send_email | Human approval, idempotency key |
| Administrative | delete_record, rotate_key | Not exposed to agents. At all. |
The fourth row is the one people argue about. Our position: if the blast radius of a mistake is "restore from backup", the tool does not go behind a probabilistic caller. Build a deterministic workflow with an agent-triggered approval step instead. You lose very little — the agent still initiates, a human still clicks — and you remove an entire category of postmortem.
For the third tier, idempotency keys matter more than they look. Agents retry. A network blip during refund_order becomes two refunds unless the tool call carries a key derived from the trajectory step, not generated fresh per attempt.
Making it observable
An agent calling tools through a gateway produces exactly the shape of data that distributed tracing was built for, so use it. OpenTelemetry's GenAI semantic conventions are the vendor-neutral standard now, and instrumenting once against them beats instrumenting per vendor.
The span tree we emit per turn:
turn (gen_ai.operation.name=chat)
├── model.completion gen_ai.request.model, gen_ai.usage.input_tokens
├── tool.select n_candidates, n_selected
├── tool.call gen_ai.tool.name, tier, policy.decision, latency
│ └── downstream.http
├── tool.call …
└── model.completion (final)
What you want from this, in order of how often it saves you:
- Cost per turn, broken down by tool. The expensive agent is almost never expensive because of the model. It is expensive because one tool returns 40KB of JSON that goes straight back into context.
- Tool call distribution. A tool that is never selected is dead weight in the catalogue. A tool selected 80% of the time probably wants to be a deterministic step.
- Policy denials. A rising denial rate means the agent is repeatedly trying something it cannot do, which usually means the prompt promises a capability the policy forbids.
- Trajectory depth. Turns that take eleven tool calls when the median is three are your failure population. Sample them into an eval set.
What we actually ship
For a team putting agents in front of internal users for the first time, the reference architecture is deliberately boring:
- One gateway, in the request path, holding all credentials. Not a library — a service, because it needs to be the enforcement point even when someone writes a new agent runtime.
- Token exchange per tool call, audience-bound, short-lived.
- A tool registry with tier, owner and description reviewed as a set. Registration is a pull request, not a config reload.
- Static scoping always; retrieval-based selection only past thirty tools.
- OTel spans on every tool call, with cost and policy decision as attributes.
- Irreversible operations behind an approval step, with idempotency keys derived from the trajectory.
Things we deliberately leave out of a first deployment: a bespoke MCP server for every internal system (wrap the two that matter, and use the API directly for the rest until the pattern proves itself), and any attempt at a universal permission model. Start with per-agent allowlists. They are ugly, they are explicit, and they are auditable — which is the only property that matters in the review where someone asks what this thing can reach.
The part that will change
MCP's 2026 roadmap is pointed squarely at this: transport scalability, governance, OAuth 2.1 and audit trails as first-class concerns rather than things every team rebuilds. Some of what is described above will move into the protocol and the reference servers, and that will be a good day.
Until then it is your gateway's job. Build it as if it will be the most security-relevant service you own, because for a while it will be.
The bar is not "the agent can call our tools." Any weekend gets you that. The bar is that a compliance reviewer can ask which humans, through which agent, touched which records last Tuesday — and you can answer it from one index, in one query, without qualifying the answer.