github / github/copilot-cli

Concurrent MCP OAuth token-refresh for two servers cancels one reconnect (self-heals, but surfaces a false hard-failure error)

Đang mở
#4,842 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

triage
Ngôn ngữ chính
Shell
Star
11.2k
Fork
1.9k
Merge trung bình
14 giờ 16 phút
Pull request đã merge (30 ngày)
6

Mô tả

Describe the bug

When two remote HTTP MCP servers (atlassian-mcp and ado-remote-mcp) both receive a 401 OAuth
challenge at nearly the same moment, the concurrent background token-refresh/reconnect for one of
them gets cancelled (quit_reason":"Cancelled"), producing a foreground-facing error:

Failed to refresh foreground session after MCP OAuth: Error: MCP server "ado-remote-mcp" failed to
reconnect: MCP server "ado-remote-mcp" connection was cancelled

The CLI self-heals a few seconds later (Successfully authenticated with ado-remote-mcp), so the
server ends up usable, but the error surfaces to the user as an alarming failure
(Failed to connect to MCP server "ado-remote-mcp": ... connection was cancelled. Execute '/mcp show ado-remote-mcp' to inspect or check the logs.) for a condition that isn't actually a lasting failure.

This looks related to (but distinct from) #4753, #4084, and #3706 — those involve session-resume
handover, OAuth routed to the wrong handler, and reconnect fan-out across many hosts, respectively.
This report is specifically about two servers hitting an expired-token 401 at the same time,
where the reconnect task for one server is cancelled — apparently pre-empted by the other server's
concurrent OAuth flow — and only recovers via a fully independent retry moments later.

Affected version

1.0.83

Steps to reproduce
  1. Configure two remote HTTP MCP servers with OAuth (e.g. atlassian-mcp and an Azure DevOps
    remote MCP server) such that both access tokens expire around the same time.
  2. Start (or resume) a session after both tokens have expired.
  3. Both servers receive a 401 at nearly the same timestamp and each begins a background
    OAuth refresh/reconnect.
  4. One server's reconnect task is cancelled mid-flight; the CLI logs an ERROR and (in this
    instance) also surfaces Failed to refresh foreground session after MCP OAuth: ... connection was cancelled to the user.
  5. ~3 seconds later, a fresh authentication attempt for the same server succeeds
    (Successfully authenticated with <server>), and the server becomes usable — but only after
    the alarming error has already been shown.
Actual log excerpt (redacted server URLs)
23:12:43.336Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
  when Client(OAuthChallenge { ... "https://mcp.atlassian.com/.well-known/oauth-protected-resource/v2/mcp" ...
  response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:43.535Z [ERROR] Refreshing authentication for atlassian-mcp...
23:12:43.972Z [ERROR] Successfully authenticated with atlassian-mcp
23:12:44.106Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
  when Client(OAuthChallenge { ... "https://mcp.dev.azure.com/.well-known/oauth-protected-resource/<org>" ...
  response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:45.148Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed, ... (atlassian, again)
23:12:45.327Z [INFO]  [rust:rmcp::service] task cancelled
23:12:45.327Z [INFO]  [rust:rmcp::service] task cancelled
23:12:45.327Z [WARNING] [rust:copilot_runtime::session::mcp::agent_host] MCP background reconnect after
  OAuth still needs auth {"server":"ado-remote-mcp"}
23:12:45.327Z [INFO]  [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
23:12:45.752Z [ERROR] Failed to refresh foreground session after MCP OAuth: Error: MCP server
  "ado-remote-mcp" failed to reconnect: MCP server "ado-remote-mcp" connection was cancelled
23:12:48.220Z [ERROR] Successfully authenticated with ado-remote-mcp
Expected behavior
  • Concurrent token refreshes for independent MCP servers should not cancel one another's
    reconnect task.
  • If a reconnect is transiently cancelled but recovers within a few seconds, that should not be
    surfaced to the user as a hard failure (or should be labeled as transient/retrying, not as an
    actionable error pointing at /mcp show).
  • Ideally the retry/backoff for concurrent OAuth refreshes across multiple servers should be
    serialized or otherwise made non-cancelling of sibling in-flight reconnects.
Environment
  • OS: Windows 11 (build 26200)
  • CLI version: 1.0.83 (exe built 2026-09-04)
  • MCP servers involved: atlassian-mcp (remote HTTP, OAuth), ado-remote-mcp (remote HTTP, OAuth,
    Azure DevOps)
  • Both servers configured as type: "http" with OAuth in .copilot/mcp-config.json / repo
    .github/mcp.json
Additional context

Ruled out as a cause: duplicate/conflicting MCP server definitions across repo (.github/mcp.json)
and local (~/.copilot/mcp-config.json) scopes for the same server name — the cancellation occurs
during the concurrent-401 OAuth refresh path itself (rmcp::service/agent_host logs), independent
of which scope's config won. Removing the local duplicate did not change this behavior on retest;
it's a config hygiene fix but not the root cause of the cancellation.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Hướng nghiên cứu

Bắt đầu bằng cách lần theo luồng kết nối lại OAuth đồng thời thông qua các log có tên copilot_runtime::session::mcp::agent_host và rmcp::service, tập trung vào việc hủy và truyền lỗi lên tiến trình foreground. Tái hiện hai phản hồi 401 đồng thời và xác minh rằng việc kết nối lại của một máy chủ không hủy việc kết nối lại của máy chủ kia, đồng thời việc khôi phục tạm thời không bị báo cáo là lỗi nghiêm trọng.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
rust
Lĩnh vực
authentication, cli
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Sôi nổi
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
48/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.