Concurrent MCP OAuth token-refresh for two servers cancels one reconnect (self-heals, but surfaces a false hard-failure error)
Personne n'a encore pris cette issue.
- Langage dominant
- Shell
- Étoiles
- 11.2k
- Forks
- 1.9k
- Merge moyen
- 14 h 16 min
- PR mergées (30 j)
- 6
Description
Describe the bug
When two remote HTTP MCP servers (atlassian-mcp and ado-remote-mcp) both receive a 401 OAuth
challenge at nearly the same moment, the concurrent background token-refresh/reconnect for one of
them gets cancelled (quit_reason":"Cancelled"), producing a foreground-facing error:
Failed to refresh foreground session after MCP OAuth: Error: MCP server "ado-remote-mcp" failed to
reconnect: MCP server "ado-remote-mcp" connection was cancelled
The CLI self-heals a few seconds later (Successfully authenticated with ado-remote-mcp), so the
server ends up usable, but the error surfaces to the user as an alarming failure
(Failed to connect to MCP server "ado-remote-mcp": ... connection was cancelled. Execute '/mcp show ado-remote-mcp' to inspect or check the logs.) for a condition that isn't actually a lasting failure.
This looks related to (but distinct from) #4753, #4084, and #3706 — those involve session-resume
handover, OAuth routed to the wrong handler, and reconnect fan-out across many hosts, respectively.
This report is specifically about two servers hitting an expired-token 401 at the same time,
where the reconnect task for one server is cancelled — apparently pre-empted by the other server's
concurrent OAuth flow — and only recovers via a fully independent retry moments later.
Affected version
1.0.83
Steps to reproduce
- Configure two remote HTTP MCP servers with OAuth (e.g.
atlassian-mcpand an Azure DevOps
remote MCP server) such that both access tokens expire around the same time. - Start (or resume) a session after both tokens have expired.
- Both servers receive a
401at nearly the same timestamp and each begins a background
OAuth refresh/reconnect. - One server's reconnect task is cancelled mid-flight; the CLI logs an
ERRORand (in this
instance) also surfacesFailed to refresh foreground session after MCP OAuth: ... connection was cancelledto the user. - ~3 seconds later, a fresh authentication attempt for the same server succeeds
(Successfully authenticated with <server>), and the server becomes usable — but only after
the alarming error has already been shown.
Actual log excerpt (redacted server URLs)
23:12:43.336Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
when Client(OAuthChallenge { ... "https://mcp.atlassian.com/.well-known/oauth-protected-resource/v2/mcp" ...
response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:43.535Z [ERROR] Refreshing authentication for atlassian-mcp...
23:12:43.972Z [ERROR] Successfully authenticated with atlassian-mcp
23:12:44.106Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
when Client(OAuthChallenge { ... "https://mcp.dev.azure.com/.well-known/oauth-protected-resource/<org>" ...
response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:45.148Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed, ... (atlassian, again)
23:12:45.327Z [INFO] [rust:rmcp::service] task cancelled
23:12:45.327Z [INFO] [rust:rmcp::service] task cancelled
23:12:45.327Z [WARNING] [rust:copilot_runtime::session::mcp::agent_host] MCP background reconnect after
OAuth still needs auth {"server":"ado-remote-mcp"}
23:12:45.327Z [INFO] [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
23:12:45.752Z [ERROR] Failed to refresh foreground session after MCP OAuth: Error: MCP server
"ado-remote-mcp" failed to reconnect: MCP server "ado-remote-mcp" connection was cancelled
23:12:48.220Z [ERROR] Successfully authenticated with ado-remote-mcp
Expected behavior
- Concurrent token refreshes for independent MCP servers should not cancel one another's
reconnect task. - If a reconnect is transiently cancelled but recovers within a few seconds, that should not be
surfaced to the user as a hard failure (or should be labeled as transient/retrying, not as an
actionable error pointing at/mcp show). - Ideally the retry/backoff for concurrent OAuth refreshes across multiple servers should be
serialized or otherwise made non-cancelling of sibling in-flight reconnects.
Environment
- OS: Windows 11 (build 26200)
- CLI version: 1.0.83 (exe built 2026-09-04)
- MCP servers involved:
atlassian-mcp(remote HTTP, OAuth),ado-remote-mcp(remote HTTP, OAuth,
Azure DevOps) - Both servers configured as
type: "http"with OAuth in.copilot/mcp-config.json/ repo
.github/mcp.json
Additional context
Ruled out as a cause: duplicate/conflicting MCP server definitions across repo (.github/mcp.json)
and local (~/.copilot/mcp-config.json) scopes for the same server name — the cancellation occurs
during the concurrent-401 OAuth refresh path itself (rmcp::service/agent_host logs), independent
of which scope's config won. Removing the local duplicate did not change this behavior on retest;
it's a config hygiene fix but not the root cause of the cancellation.
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Piste de recherche
Commencez par suivre le chemin de reconnexion OAuth concurrente à travers les logs nommés copilot_runtime::session::mcp::agent_host et rmcp::service, en vous concentrant sur l’annulation et la propagation des erreurs vers l’avant-plan. Reproduisez deux réponses 401 simultanées et vérifiez que la reconnexion d’un serveur n’annule pas celle de l’autre, et qu’une récupération transitoire n’est pas signalée comme un échec critique.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- rust
- Domaine
- authentication, cli
- Type d'issue
- Bug
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- Active
- Clarté
- Plutôt claire
- Accessibilité débutants
- 48/100