Concurrent MCP OAuth token-refresh for two servers cancels one reconnect (self-heals, but surfaces a false hard-failure error)
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Shell
- Sterne
- 11.2k
- Forks
- 1.9k
- Ø Merge
- 14 Std. 16 Min.
- Gemergte PRs (30 T.)
- 6
Beschreibung
Describe the bug
When two remote HTTP MCP servers (atlassian-mcp and ado-remote-mcp) both receive a 401 OAuth
challenge at nearly the same moment, the concurrent background token-refresh/reconnect for one of
them gets cancelled (quit_reason":"Cancelled"), producing a foreground-facing error:
Failed to refresh foreground session after MCP OAuth: Error: MCP server "ado-remote-mcp" failed to
reconnect: MCP server "ado-remote-mcp" connection was cancelled
The CLI self-heals a few seconds later (Successfully authenticated with ado-remote-mcp), so the
server ends up usable, but the error surfaces to the user as an alarming failure
(Failed to connect to MCP server "ado-remote-mcp": ... connection was cancelled. Execute '/mcp show ado-remote-mcp' to inspect or check the logs.) for a condition that isn't actually a lasting failure.
This looks related to (but distinct from) #4753, #4084, and #3706 — those involve session-resume
handover, OAuth routed to the wrong handler, and reconnect fan-out across many hosts, respectively.
This report is specifically about two servers hitting an expired-token 401 at the same time,
where the reconnect task for one server is cancelled — apparently pre-empted by the other server's
concurrent OAuth flow — and only recovers via a fully independent retry moments later.
Affected version
1.0.83
Steps to reproduce
- Configure two remote HTTP MCP servers with OAuth (e.g.
atlassian-mcpand an Azure DevOps
remote MCP server) such that both access tokens expire around the same time. - Start (or resume) a session after both tokens have expired.
- Both servers receive a
401at nearly the same timestamp and each begins a background
OAuth refresh/reconnect. - One server's reconnect task is cancelled mid-flight; the CLI logs an
ERRORand (in this
instance) also surfacesFailed to refresh foreground session after MCP OAuth: ... connection was cancelledto the user. - ~3 seconds later, a fresh authentication attempt for the same server succeeds
(Successfully authenticated with <server>), and the server becomes usable — but only after
the alarming error has already been shown.
Actual log excerpt (redacted server URLs)
23:12:43.336Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
when Client(OAuthChallenge { ... "https://mcp.atlassian.com/.well-known/oauth-protected-resource/v2/mcp" ...
response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:43.535Z [ERROR] Refreshing authentication for atlassian-mcp...
23:12:43.972Z [ERROR] Successfully authenticated with atlassian-mcp
23:12:44.106Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
when Client(OAuthChallenge { ... "https://mcp.dev.azure.com/.well-known/oauth-protected-resource/<org>" ...
response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:45.148Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed, ... (atlassian, again)
23:12:45.327Z [INFO] [rust:rmcp::service] task cancelled
23:12:45.327Z [INFO] [rust:rmcp::service] task cancelled
23:12:45.327Z [WARNING] [rust:copilot_runtime::session::mcp::agent_host] MCP background reconnect after
OAuth still needs auth {"server":"ado-remote-mcp"}
23:12:45.327Z [INFO] [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
23:12:45.752Z [ERROR] Failed to refresh foreground session after MCP OAuth: Error: MCP server
"ado-remote-mcp" failed to reconnect: MCP server "ado-remote-mcp" connection was cancelled
23:12:48.220Z [ERROR] Successfully authenticated with ado-remote-mcp
Expected behavior
- Concurrent token refreshes for independent MCP servers should not cancel one another's
reconnect task. - If a reconnect is transiently cancelled but recovers within a few seconds, that should not be
surfaced to the user as a hard failure (or should be labeled as transient/retrying, not as an
actionable error pointing at/mcp show). - Ideally the retry/backoff for concurrent OAuth refreshes across multiple servers should be
serialized or otherwise made non-cancelling of sibling in-flight reconnects.
Environment
- OS: Windows 11 (build 26200)
- CLI version: 1.0.83 (exe built 2026-09-04)
- MCP servers involved:
atlassian-mcp(remote HTTP, OAuth),ado-remote-mcp(remote HTTP, OAuth,
Azure DevOps) - Both servers configured as
type: "http"with OAuth in.copilot/mcp-config.json/ repo
.github/mcp.json
Additional context
Ruled out as a cause: duplicate/conflicting MCP server definitions across repo (.github/mcp.json)
and local (~/.copilot/mcp-config.json) scopes for the same server name — the cancellation occurs
during the concurrent-401 OAuth refresh path itself (rmcp::service/agent_host logs), independent
of which scope's config won. Removing the local duplicate did not change this behavior on retest;
it's a config hygiene fix but not the root cause of the cancellation.
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne damit, den Pfad für die gleichzeitige OAuth-Wiederverbindung anhand der benannten copilot_runtime::session::mcp::agent_host- und rmcp::service-Logs nachzuverfolgen, wobei der Schwerpunkt auf der Abbruchbehandlung und der Weitergabe von Fehlern an den Vordergrund liegt. Reproduziere zwei gleichzeitige 401-Antworten und überprüfe, dass die Wiederverbindung eines Servers nicht die des anderen abbricht und dass eine vorübergehende Wiederherstellung nicht als schwerwiegender Fehler gemeldet wird.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- rust
- Bereich
- authentication, cli
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Aktiv
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 48/100