Skip to content

MCP Token Passthrough: The Shortcut That Turns Your Server Into a Confused Deputy

6 min read Fullmakt Team

  • mcp
  • authentication
  • credentials
  • traceability
  • governance
  • agents
  • security

An MCP server needs to call your CRM on behalf of an AI agent. The agent already arrived with a valid access token, so the quickest implementation is to take the Authorization header from the incoming request and attach it to the outgoing one. Two lines of code, no extra login, it works on the first try.

That shortcut is called token passthrough, and the MCP authorization specification explicitly forbids it. It is also one of the most common ways an agentic AI deployment ends up with a login handshake that looks secure on the diagram and isn’t in practice.

What token passthrough actually is

In a correct MCP flow the client obtains a token for the MCP server. The server validates it, then does its downstream work using a separate credential it obtained for that downstream service. Two tokens, two audiences.

With passthrough there is one token and it travels the whole chain:

  1. The agent sends token T to the MCP server.
  2. The server doesn’t check who T was issued for, or doesn’t care.
  3. The server forwards T unchanged to the CRM, the ticketing system or the payments API.

Everything downstream now trusts a token that was never minted for it, and the MCP server has quietly become a proxy for credentials it doesn’t own.

Four things that break

Audience binding disappears. OAuth access tokens carry an audience: the resource they’re valid for. RFC 8707 resource indicators let a client say “this token is for https://mcp.example.com and nothing else”. A server that forwards the token ignores that binding. A token meant for one service is now accepted by another, which is the textbook confused deputy.

Blast radius multiplies. If T leaks from the MCP server’s logs, memory or a prompt-injected tool response, the attacker holds a working credential for every system the server ever forwarded it to, not just the MCP server.

Downstream controls see the wrong caller. Rate limits, scopes and policy on the downstream API apply to whoever the token says is calling. When many agents and tools share that path, the downstream system can’t tell them apart, and it can’t apply least privilege per agent.

Traceability collapses. The downstream audit log says “user X called the API”. It cannot say that an agent called it, through which MCP server, inside which delegation chain. When something goes wrong, incident forensics starts from a log that has already lost the facts you need.

A realistic failure

Picture an internal “meeting notes” MCP server that also calls a finance API using passthrough. A poisoned calendar invite contains instructions telling the agent to call a tool with a crafted argument. The tool handler logs the full request for debugging, headers included. An engineer pastes that log into a ticket. The token in it was issued to the agent with broad scopes and a one-hour lifetime. Nothing was “hacked”: a debug log became a credential leak, and the token worked against the finance API because passthrough meant nobody checked the audience.

What good looks like

The fix is boring and well understood. Treat each hop as its own trust boundary:

  • Validate the audience on every inbound token. Reject tokens whose aud is not this server. Compare the value exactly and canonically so trailing-slash and case differences can’t open a gap.
  • Bind tokens with resource indicators. The client sends resource on both the authorize and token requests, and the issuer refuses a code exchange whose resource differs from the one the code was issued for.
  • Use PKCE with S256. It protects the code exchange itself from interception.
  • Mint downstream credentials separately. Use token exchange (RFC 8693) or a broker that holds the downstream secret, scoped to one task and short-lived.
  • Keep tokens off surfaces they weren’t issued for. An agent token should be refused outright on endpoints outside its remit.
  • Log at the action boundary. Record agent identity, tool, target and outcome, never the token itself.

How Fullmakt approaches it

Fullmakt’s MCP endpoint follows this pattern. Its OAuth surface is authorization-code with PKCE (S256 only) and advertises its protected resource metadata per RFC 9728. When a client sends an RFC 8707 resource, it must name this server’s MCP endpoint, and the token request must match the resource the authorization code was issued for. The issued connector token carries that audience, and refresh tokens rotate. Agent tokens used on the wrong surface get a 403 wrong_token_surface rather than being quietly accepted.

Downstream, the model never holds the upstream secret. Fullmakt resolves a credential reference server-side at call time, applies policy, and records an audit entry per request, so the log answers which agent made which call. See scoped credentials for AI agents and auditing every agent API call for the details.

To be precise about scope: Fullmakt secures the path between your agents and the APIs they call. It can’t fix an MCP server you run yourself that forwards tokens. That server needs the audience check above.

A quick audit you can run this week

  1. For each MCP server you operate, find where it builds outbound requests. If it copies the inbound Authorization header, you have passthrough.
  2. Decode a token from each client and check aud. If it names a service the token shouldn’t reach, fix issuance.
  3. Grep logs for bearer tokens. Redact headers by default.
  4. Confirm a downstream audit entry names the agent, not only the human.
  5. Shorten access-token lifetimes and rotate refresh tokens.

The business case

Passthrough is cheap to build and expensive to explain to an auditor. When a security review asks “which credential did this agent use against the finance system, and who approved it?”, a separate scoped credential and a per-call audit record give a one-line answer. That shortens vendor security reviews, supports SOC 2 and ISO 42001 evidence, and keeps one leaked token from becoming an incident across every connected system.

Running MCP servers or agents against production APIs? Try Fullmakt and keep credentials out of the model and out of the passthrough chain.

FAQ

What is token passthrough in MCP? It is when an MCP server forwards the access token it received from a client to a downstream API instead of obtaining a separate credential for that API. The MCP authorization specification prohibits it.

Why is token passthrough dangerous? It bypasses audience binding, lets one stolen token reach many systems, prevents downstream systems from identifying the real caller, and removes the agent from the audit trail.

What is an OAuth audience and why does it matter for AI agents? The audience (aud) names the service a token is valid for. Checking it stops a token issued for one service from being accepted by another, which is the core defence against confused-deputy attacks.

What are RFC 8707 resource indicators? A resource parameter that lets a client request a token for a specific server. The issuer binds the token to that resource and the server verifies it.

How should an MCP server call downstream APIs? With its own credential for that API: obtained via token exchange, or held and injected by a credential broker, scoped to the task and short-lived.

Does Fullmakt prevent token passthrough? For its own MCP endpoint, yes: it enforces audience and resource binding and keeps upstream credentials server-side. It can’t change how a third-party or self-built MCP server handles tokens.