Concurrent MCP OAuth token-refresh for two servers cancels one reconnect (self-heals, but surfaces a false hard-failure error)
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Shell
- Star
- 11.2k
- Fork
- 1.9k
- Merge trung bình
- 14 giờ 16 phút
- Pull request đã merge (30 ngày)
- 6
Mô tả
Describe the bug
When two remote HTTP MCP servers (atlassian-mcp and ado-remote-mcp) both receive a 401 OAuth
challenge at nearly the same moment, the concurrent background token-refresh/reconnect for one of
them gets cancelled (quit_reason":"Cancelled"), producing a foreground-facing error:
Failed to refresh foreground session after MCP OAuth: Error: MCP server "ado-remote-mcp" failed to
reconnect: MCP server "ado-remote-mcp" connection was cancelled
The CLI self-heals a few seconds later (Successfully authenticated with ado-remote-mcp), so the
server ends up usable, but the error surfaces to the user as an alarming failure
(Failed to connect to MCP server "ado-remote-mcp": ... connection was cancelled. Execute '/mcp show ado-remote-mcp' to inspect or check the logs.) for a condition that isn't actually a lasting failure.
This looks related to (but distinct from) #4753, #4084, and #3706 — those involve session-resume
handover, OAuth routed to the wrong handler, and reconnect fan-out across many hosts, respectively.
This report is specifically about two servers hitting an expired-token 401 at the same time,
where the reconnect task for one server is cancelled — apparently pre-empted by the other server's
concurrent OAuth flow — and only recovers via a fully independent retry moments later.
Affected version
1.0.83
Steps to reproduce
- Configure two remote HTTP MCP servers with OAuth (e.g.
atlassian-mcpand an Azure DevOps
remote MCP server) such that both access tokens expire around the same time. - Start (or resume) a session after both tokens have expired.
- Both servers receive a
401at nearly the same timestamp and each begins a background
OAuth refresh/reconnect. - One server's reconnect task is cancelled mid-flight; the CLI logs an
ERRORand (in this
instance) also surfacesFailed to refresh foreground session after MCP OAuth: ... connection was cancelledto the user. - ~3 seconds later, a fresh authentication attempt for the same server succeeds
(Successfully authenticated with <server>), and the server becomes usable — but only after
the alarming error has already been shown.
Actual log excerpt (redacted server URLs)
23:12:43.336Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
when Client(OAuthChallenge { ... "https://mcp.atlassian.com/.well-known/oauth-protected-resource/v2/mcp" ...
response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:43.535Z [ERROR] Refreshing authentication for atlassian-mcp...
23:12:43.972Z [ERROR] Successfully authenticated with atlassian-mcp
23:12:44.106Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed,
when Client(OAuthChallenge { ... "https://mcp.dev.azure.com/.well-known/oauth-protected-resource/<org>" ...
response: McpOAuthHttpResponse { status_code: 401, ... } })
23:12:45.148Z [ERROR] [rust:rmcp::transport::worker] worker quit with fatal: Transport channel closed, ... (atlassian, again)
23:12:45.327Z [INFO] [rust:rmcp::service] task cancelled
23:12:45.327Z [INFO] [rust:rmcp::service] task cancelled
23:12:45.327Z [WARNING] [rust:copilot_runtime::session::mcp::agent_host] MCP background reconnect after
OAuth still needs auth {"server":"ado-remote-mcp"}
23:12:45.327Z [INFO] [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
23:12:45.752Z [ERROR] Failed to refresh foreground session after MCP OAuth: Error: MCP server
"ado-remote-mcp" failed to reconnect: MCP server "ado-remote-mcp" connection was cancelled
23:12:48.220Z [ERROR] Successfully authenticated with ado-remote-mcp
Expected behavior
- Concurrent token refreshes for independent MCP servers should not cancel one another's
reconnect task. - If a reconnect is transiently cancelled but recovers within a few seconds, that should not be
surfaced to the user as a hard failure (or should be labeled as transient/retrying, not as an
actionable error pointing at/mcp show). - Ideally the retry/backoff for concurrent OAuth refreshes across multiple servers should be
serialized or otherwise made non-cancelling of sibling in-flight reconnects.
Environment
- OS: Windows 11 (build 26200)
- CLI version: 1.0.83 (exe built 2026-09-04)
- MCP servers involved:
atlassian-mcp(remote HTTP, OAuth),ado-remote-mcp(remote HTTP, OAuth,
Azure DevOps) - Both servers configured as
type: "http"with OAuth in.copilot/mcp-config.json/ repo
.github/mcp.json
Additional context
Ruled out as a cause: duplicate/conflicting MCP server definitions across repo (.github/mcp.json)
and local (~/.copilot/mcp-config.json) scopes for the same server name — the cancellation occurs
during the concurrent-401 OAuth refresh path itself (rmcp::service/agent_host logs), independent
of which scope's config won. Removing the local duplicate did not change this behavior on retest;
it's a config hygiene fix but not the root cause of the cancellation.
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Bắt đầu bằng cách lần theo luồng kết nối lại OAuth đồng thời thông qua các log có tên copilot_runtime::session::mcp::agent_host và rmcp::service, tập trung vào việc hủy và truyền lỗi lên tiến trình foreground. Tái hiện hai phản hồi 401 đồng thời và xác minh rằng việc kết nối lại của một máy chủ không hủy việc kết nối lại của máy chủ kia, đồng thời việc khôi phục tạm thời không bị báo cáo là lỗi nghiêm trọng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- rust
- Lĩnh vực
- authentication, cli
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Sôi nổi
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 48/100