github / github/copilot-cli

Concurrent MCP OAuth token-refresh for two servers cancels one reconnect (self-heals, but surfaces a false hard-failure error)

Ouverte
#4,842 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

triage
Langage dominant
Shell
Étoiles
11.2k
Forks
1.9k
Merge moyen
14 h 16 min
PR mergées (30 j)
6

Description

Describe the bug

When two remote HTTP MCP servers (atlassian-mcp and ado-remote-mcp) both receive a 401 OAuth
challenge at nearly the same moment, the concurrent background token-refresh/reconnect for one of
them gets cancelled (quit_reason":"Cancelled"), producing a foreground-facing error:

Failed to refresh foreground session after MCP OAuth: Error: MCP server "ado-remote-mcp" failed to
reconnect: MCP server "ado-remote-mcp" connection was cancelled

The CLI self-heals a few seconds later (Successfully authenticated with ado-remote-mcp), so the
server ends up usable, but the error surfaces to the user as an alarming failure
(Failed to connect to MCP server "ado-remote-mcp": ... connection was cancelled. Execute '/mcp show ado-remote-mcp' to inspect or check the logs.) for a condition that isn't actually a lasting failure.

This looks related to (but distinct from) #4753, #4084, and #3706 — those involve session-resume
handover, OAuth routed to the wrong handler, and reconnect fan-out across many hosts, respectively.
This report is specifically about two servers hitting an expired-token 401 at the same time,
where the reconnect task for one server is cancelled — apparently pre-empted by the other server's
concurrent OAuth flow — and only recovers via a fully independent retry moments later.

Affected version

1.0.83

Steps to reproduce
  1. Configure two remote HTTP MCP servers with OAuth (e.g. atlassian-mcp and an Azure DevOps
    remote MCP server) such that both access tokens expire around the same time.
  2. Start (or resume) a session after both tokens have expired.
  3. Both servers receive a 401 at nearly the same timestamp and each begins a background
    OAuth refresh/reconnect.
  4. One server's reconnect task is cancelled mid-flight; the CLI logs an ERROR and (in this
    instance) also surfaces Failed to refresh foreground session after MCP OAuth: ... connection was cancelled to the user.
  5. ~3 seconds later, a fresh authentication attempt for the same server succeeds
    (Successfully authenticated with <server>), and the server becomes usable — but only after
    the alarming error has already been shown.
Actual log excerpt (redacted server URLs)
23:12:43.336Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
  when Client(OAuthChallenge { ... "https://mcp.atlassian.com/.well-known/oauth-protected-resource/v2/mcp" ...
  response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:43.535Z [ERROR] Refreshing authentication for atlassian-mcp...
23:12:43.972Z [ERROR] Successfully authenticated with atlassian-mcp
23:12:44.106Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
  when Client(OAuthChallenge { ... "https://mcp.dev.azure.com/.well-known/oauth-protected-resource/<org>" ...
  response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:45.148Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed, ... (atlassian, again)
23:12:45.327Z [INFO]  [rust:rmcp::service] task cancelled
23:12:45.327Z [INFO]  [rust:rmcp::service] task cancelled
23:12:45.327Z [WARNING] [rust:copilot_runtime::session::mcp::agent_host] MCP background reconnect after
  OAuth still needs auth {"server":"ado-remote-mcp"}
23:12:45.327Z [INFO]  [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
23:12:45.752Z [ERROR] Failed to refresh foreground session after MCP OAuth: Error: MCP server
  "ado-remote-mcp" failed to reconnect: MCP server "ado-remote-mcp" connection was cancelled
23:12:48.220Z [ERROR] Successfully authenticated with ado-remote-mcp
Expected behavior
  • Concurrent token refreshes for independent MCP servers should not cancel one another's
    reconnect task.
  • If a reconnect is transiently cancelled but recovers within a few seconds, that should not be
    surfaced to the user as a hard failure (or should be labeled as transient/retrying, not as an
    actionable error pointing at /mcp show).
  • Ideally the retry/backoff for concurrent OAuth refreshes across multiple servers should be
    serialized or otherwise made non-cancelling of sibling in-flight reconnects.
Environment
  • OS: Windows 11 (build 26200)
  • CLI version: 1.0.83 (exe built 2026-09-04)
  • MCP servers involved: atlassian-mcp (remote HTTP, OAuth), ado-remote-mcp (remote HTTP, OAuth,
    Azure DevOps)
  • Both servers configured as type: "http" with OAuth in .copilot/mcp-config.json / repo
    .github/mcp.json
Additional context

Ruled out as a cause: duplicate/conflicting MCP server definitions across repo (.github/mcp.json)
and local (~/.copilot/mcp-config.json) scopes for the same server name — the cancellation occurs
during the concurrent-401 OAuth refresh path itself (rmcp::service/agent_host logs), independent
of which scope's config won. Removing the local duplicate did not change this behavior on retest;
it's a config hygiene fix but not the root cause of the cancellation.

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez par suivre le chemin de reconnexion OAuth concurrente à travers les logs nommés copilot_runtime::session::mcp::agent_host et rmcp::service, en vous concentrant sur l’annulation et la propagation des erreurs vers l’avant-plan. Reproduisez deux réponses 401 simultanées et vérifiez que la reconnexion d’un serveur n’annule pas celle de l’autre, et qu’une récupération transitoire n’est pas signalée comme un échec critique.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
rust
Domaine
authentication, cli
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Active
Clarté
Plutôt claire
Accessibilité débutants
48/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.