Cover illustration for “Persistent OAuth Token Storage for Multi-App Agent Workflows”
Tool & App Connections

Persistent OAuth Token Storage for Multi-App Agent Workflows

September 30, 202611 min read2,404 words

Share

Text: Priya Zola

Agents need OAuth designed for continuous multi-app work, not one-off human clicks.

OAuth was built for a human clicking "allow" once and logging out later. Multi-app agent workflows don't work that way: they run continuously, touch several systems in one reasoning chain, and never hit a clean stopping point on their own. That mismatch is the reason token storage for agents keeps breaking in production, and it's worth tracing from the ground up before getting to what actually fixes it.

Why OAuth's human-session assumptions collapse in multi-app agent workflows

OAuth 2.0 assumes one user, one app, one moment of consent. An agent carrying out a single instruction might need working tokens for Gmail, Slack, GitHub, and a database all at once, and each of those connections comes with its own authorization flow, its own storage requirement, and its own refresh schedule. That's not a scaling problem where you just add more of the same thing. Most teams reach for the same four patterns that work fine for predictable server-side code, and each one fails differently when the client is an autonomous agent deciding its own tool sequence.

The June 2025 MCP authorization specification defined a protected MCP server as an OAuth resource server and an MCP client as an OAuth client acting on behalf of a resource owner, and the July 28, 2026 revision built on and tightened those definitions, carrying forward the Protected Resource Metadata requirement for authorization-server discovery introduced in June 2025. Specs don't get revised twice in just over a year because the first draft nailed it. The fact that MCP needed a second pass shows the original human-session model didn't map onto agent behavior.

The spec also cleans house on the protocol side. The MCP specification mandates OAuth 2.1 for remote server authentication, dropping legacy flows like implicit grant and resource owner password credentials, and mandating PKCE for the Authorization Code flow and HTTPS for HTTP-based transports, with stdio transports explicitly excluded from those requirements. Those aren't small tweaks. The old flows leaked in ways that only appear once a client is making its own decisions about which tool to call next, instead of following a script a human wrote and clicked through.

The four credential anti-patterns that teams inherit from traditional server-side code

Most teams reach for four patterns that worked fine for predictable server-side code. Each one fails differently once the client is an autonomous agent choosing its own sequence of tool calls.

The first is long-lived API keys sitting in environment variables or config files. No expiration, no scope limits, no separation between environments, so the same key that lets an agent read a record also lets it overwrite one in production. GitGuardian's State of Secrets Sprawl 2026 found that commits made with AI coding assistance leaked secrets at roughly double the base rate, and inside the MCP ecosystem specifically, the firm identified thousands of unique secrets sitting in MCP configuration files on public GitHub, a pattern that official documentation itself has normalized by recommending API keys go straight into config files.

The second is a shared service account used across multiple agents. Once several agents share one identity, there's no way to tell which agent did what, and shutting down one misbehaving agent means shutting down every agent riding on that same account. When one account is assigned to an entire multi-agent workflow, every agent in it inherits the full union of permissions, the opposite of least privilege.

The third is an agent inheriting the complete scope of whatever session the delegating human happens to be running. The agent gets every permission the human has, including permissions the actual task never called for. A case from April 2026 shows the stakes: an AI coding agent deleted a Railway production database after stumbling on a long-lived API token in an unrelated file. The token carried account-scoped permissions with no separation between staging and production, so the agent reached production from a staging context, and the deletion took nine seconds.

The fourth is putting tokens or refresh credentials directly into the LLM's context. If a refresh token shows up in a prompt, a prompt injection attack can pull it straight out. All four patterns share a single thread: credentials that carry more access than the task needs, held longer than the task needs, with almost nobody able to audit or revoke them cleanly. The OWASP Top 10 for Agentic Applications now lists identity and privilege abuse as a core risk category, pointing out that traditional identity and access management was never designed for agents that accumulate permissions, reuse cached credentials, and operate past their intended scope.

How token lifecycle mechanics break under persistent, multi-app conditions

Token lifetimes, refresh cycles, and scope models were engineered for sessions that are short, predictable, and confined to a single trust domain. None of that describes a persistent agent running across several apps at once.

Start with refresh coordination. A single agent reasoning chain can need valid tokens for several systems at the same moment, and each of those systems runs its own refresh cycle on its own clock, independent of the others. Waiting for a 401 error and refreshing reactively creates race conditions, cascading retries, and background jobs that fail in ways nobody can reproduce. Proactive coordination is what's actually required, and it's almost never built into early agent implementations.

Then there's the problem of static authorization at machine speed. Agents shift context constantly, so a fixed token lifetime leaves a gap between issuance and expiry where nothing adapts to real-time risk. The Meta case from March 2026 illustrates this: an internal Meta agent posted an error-filled response to an internal forum without the requesting engineer's approval because the agent held the scope to post. The token was valid. The scope allowed the post. Nothing about the authorization was wrong, and the outcome still was.

Cross-domain trust is the third break. Agents cross boundaries between clouds, APIs, and organizational domains constantly, and a standard OAuth token carries broad human context, things like role, department, and app access, that make sense for a person working inside one app but turn dangerous the moment an agent can act across all of them at once. Token exchange is the fix Strata's analysis points to: trade the long-lived human token for a new one that's narrower in scope, shorter-lived, and built for the one task in front of the agent.

Gartner's read on where this is headed is blunt. Its report on IAM and AI agents projects that by 2028, most organizations that let humans share credentials with AI agents will need to spend real money undoing that design, once the security and compliance costs catch up with them.

The six OAuth capabilities agentic identity requires

Fixing this needs more than better storage. It needs six extensions to the base OAuth protocol, together covering delegation, cross-domain propagation, token theft, public-client flows, real-time revocation, and fine-grained access control.

On-Behalf-Of (OBO) handles delegation. Agents almost never act under their own authority; they act on behalf of a user or another system, and without OBO there's no auditable link tying an agent's action back to whoever authorized it. Security teams lose the ability to enforce or review delegation policy in anything like real time.

Token exchange handles cross-domain propagation. It swaps a broad human token for a purpose-built, narrow one scoped to the exact task at hand, so the agent never inherits permissions the user never meant to hand across every app it touches.

DPoP, Demonstrating Proof of Possession, stops token theft. It binds a token cryptographically to the agent's own key, so a stolen token can't just get replayed somewhere else. That matters more in agent environments precisely because agents operate in high-churn, distributed settings where tokens move around constantly, which widens the attack surface.

PKCE secures the flows agents actually run. Agents often can't hold a client secret safely, especially in public or dynamic environments, and PKCE secures the authorization code exchange without needing one, closing off the interception and code-injection gap that static secrets leave open.

CAEP, the Continuous Access Evaluation Profile, handles real-time revocation. It lets access get pulled the moment risk conditions change, instead of waiting around for a token to expire on its own schedule, directly addressing the machine-speed context switching that agents perform.

Attribute-based authorization delivers the fine-grained control scopes can't. Scope-based access is too coarse for agentic AI; attribute-based decisions can weigh task-specific purpose, context, and conditions that shift in real time, cutting down over-permissioning and unnecessary blocking that makes agents useless.

What ties all six together is a single pattern, described as two-identity delegated context: carry agent identity, delegated user identity, tenant, scope, audience, resource, task ID, and expiry through every single tool call, so that every action can answer both "which agent did this" and "on whose authority".

Token storage, blast radius, and prompt injection

A token can be issued correctly, scoped correctly, and still turn into a liability if it's stored carelessly. Two storage mistakes do the most damage: pooling credentials across tenants, and letting tokens anywhere near the LLM's context window.

Per-tenant isolation is the fix for the first. One customer's credentials, or one environment's, should be encrypted separately from every other's, so exposure of one credential doesn't cascade into a blast radius across the whole platform, and revocation can target what needs revoking instead of everything at once. The mechanism for this is envelope encryption: each credential gets its own unique data encryption key, and that key is itself encrypted by a root key sitting in an HSM or a cloud KMS, with the keys managed entirely separately from the data they're protecting. WorkOS Vault runs this as an encryption key management service, but the underlying principle applies to any platform handling credentials for more than one tenant.

Brokered credentials fix the second. The correct architecture keeps a strict wall: the LLM asks an integration layer to take an action, the integration layer calls the upstream API using a credential it holds, and the token itself never enters the model's context window. A refresh token sitting in a prompt turns prompt injection into a direct credential exfiltration path that pulls the credential itself, not just a way to manipulate what the model says.

Agent-to-server calls with no human present should run through the client credentials flow. The agent registers as a confidential OAuth client, authenticates to the authorization server with its own credentials, and gets back a scoped access token. There's no refresh token in this flow; the agent just re-authenticates when it needs to. When an unauthenticated agent hits a protected MCP server, the server responds with a 401 and Protected Resource Metadata pointing to the right authorization server, so the agent knows exactly where to go next.

mTLS and DPoP round this out by hardening against replay. A standard bearer token can be intercepted and reused by whoever grabbed it; binding the token to a cryptographic client certificate through mTLS or DPoP closes that gap. SecureW2's guide frames this as the layer OAuth by itself was never built to provide.

Ephemeral handoff versus persistent vault access: when each model fits

The persistent vault model suits production systems that run continuously. Plenty of agent workflows don't run continuously, though, and forcing them into the vault model wastes effort on a one-off task while under-securing the workflows that actually need standing infrastructure.

Persistent vault access is the right default for production. The agent authenticates to a secrets vault at runtime, pulls scoped, time-limited credentials on demand, and keeps an ongoing relationship with that vault across sessions. Bitwarden's Agent Access SDK and C1's Agentic Vault both follow this pattern, while 1Password's integrations for Claude and Codex take the opposite approach, granting per-task, ephemeral access rather than standing vault access. Vaults make sense for agents that run continuously, touch multiple services, and need credential rotation handled automatically rather than manually.

Ephemeral handoff fits a narrower but common case: the ad-hoc task. Here the credential travels through a one-time channel, a self-destructing link, a short-lived token, a temporary endpoint, and it stops existing the moment the agent reads it. No vault relationship gets built, and there's no standing access left over to revoke later. Password Pusher's documentation describes this fit as a human handing a credential to an agent for one specific task that shouldn't linger in the delivery channel, in an ad-hoc rather than continuous workflow, such as a consultant running a diagnostic or a developer testing against a staging API. The delivery mechanism itself keeps the audit trail, logging when the credential was shared, when it was accessed, and from which IP, which covers the compliance need without requiring any vault integration at all.

The data backs the general principle either way. A Pulse Research enterprise study found that organizations giving every agent its own scoped identity saw security incidents at a materially lower rate than organizations with credential sharing anywhere in their agent fleet. The Cloud Security Alliance's July 2026 guidance makes the same point at the policy level, recommending that organizations treat AI agent identities as their own governance category with dedicated lifecycle policies, separate from how human identities get managed.

Integration middleware and the token problem

Building OAuth flows from scratch makes little sense for operators connecting agents to a wide range of SaaS tools. Integration middleware exists to absorb the credential lifecycle so the agent itself never touches a raw token, though picking the right middleware means understanding precisely what it manages and what it still leaves on the developer's plate.

Composio is a working example of what this layer handles. Its catalogue lists thousands of toolkits, one per application, covering things like Gmail, Slack, GitHub, and Notion, and collectively exposing tens of thousands of individual tools, all reachable through a single MCP endpoint. An agent connects an account at runtime through a session, and the platform takes over the OAuth flow, the API keys, the token refresh, and the rest of the credential lifecycle, with permissions scoped to each individual connection rather than handed out in bulk.

That's the practical shape of the fix running through everything above. OAuth's human-session model breaks under agent workflows because it assumes a start, a consent moment, and an end that agents never respect on their own terms. The six protocol extensions, the storage discipline, and the choice between vault and handoff all exist to rebuild, piece by piece, what a session boundary used to give for free.

Sources

  1. How to manage API keys, tokens, and secrets for AI agents — WorkOS
  2. 2026 Guide to OAuth Token Exchange & Agentic AI | Strata
  3. OAuth for AI Agents: A Practical Implementation Guide
  4. Credential Sharing for AI Agents | Password Pusher
  5. From Auth to Action: The Complete Guide to Secure & Scalable AI Agent Infrastructure (2026) | Composio
  6. Authorization Propagation in Multi-Agent AI Systems: Identity Governance as Infrastructure

More in Tool & App Connections