Remote/HTTP MCP servers (e.g. atlassian) are stranded 'failed' after every /clear or session relaunch
まだ誰も着手していません。
- 主要言語
- Shell
- スター
- 11.2k
- フォーク
- 1.9k
- 平均マージ
- 14時間 16分
- マージ済み PR(30日)
- 6
説明
Summary
Remote (HTTP) MCP server connections are reliably stranded in a failed state whenever the CLI performs a foreground-session handoff — which happens on every /clear, and apparently on session resume/relaunch too. The MCP connection graph is torn down and rebuilt during the handoff, and slower remote/OAuth-backed servers lose the reconnect race and never recover automatically.
Environment
- CLI version: 1.0.83 (macOS, confirmed already latest via in-app update check)
- MCP server affected:
atlassian(https://mcp.atlassian.com/v2/mcp, HTTP transport with Authorization header) - Other MCP servers configured (
github-mcp-server,codegraph,local-rag— local/stdio) are not visibly affected, likely because their reconnect is fast enough to win the race.
Steps to reproduce
- Start a Copilot CLI session with the
atlassianMCP server configured (~/.copilot/mcp-config.json, HTTP transport). - Confirm it connects fine (
/mcp show atlassian→ connected, tools listed). - Run
/clear(or relaunch/resume a session). - Run
/mcp show atlassianagain.
Expected
atlassian reconnects cleanly like the other MCP servers.
Actual
atlassian is left in:
{
"name": "atlassian",
"status": "failed",
"error": "MCP server \"atlassian\" connection was cancelled"
}
It does not self-heal — it requires a manual reconnect action or a full CLI restart, and even a full restart reproduces the same failure deterministically (confirmed across two separate relaunches, ~20 minutes apart).
Root cause (from ~/.copilot/logs/process-*.log)
Every /clear/relaunch briefly registers a throwaway foreground session, then immediately unregisters it in favor of the real one, within milliseconds:
17:27:22.425Z [INFO] Registering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.717Z [INFO] Unregistering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.720Z [INFO] Registering foreground session: d89d1d9a-99c4-4247-86fe-a75b4da5daad
17:27:28.783Z [INFO] Closing session 6266dbfb-39da-4797-9265-c37d9655a3fd
This handoff triggers a full MCP graph reload while a reload lock is still held from the previous teardown:
17:05:05.167Z [WARNING] [rust:copilot_runtime::session::mcp::session_host] MCP reload lock still held at disposal; draining the graph without it
Every MCP server — including atlassian — gets re-initialized and then almost immediately cancelled mid-handshake:
17:27:26.104Z [INFO] [rust:rmcp::service] Service initialized as client {... "name": "atlassian-mcp-server" ...}
17:27:28.809Z [INFO] [rust:rmcp::service] task cancelled
17:27:28.809Z [INFO] [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
Local/stdio servers appear to reconnect fast enough afterward to recover; atlassian's remote HTTP+OAuth handshake is slower and consistently loses the race, leaving it permanently failed with no automatic retry.
Confirmed not a network/auth issue independently:
- DNS resolves fine for
mcp.atlassian.com. curl https://mcp.atlassian.com/v2/mcpreturns401(server reachable, no valid header supplied in the manual test — expected).ATLASSIAN_MCP_AUTH_HEADERenv var is set.
Impact
Every /clear (a very common action) breaks the Atlassian MCP integration and requires a manual reconnect or CLI restart to restore it — significant daily friction for anyone using an HTTP-based MCP server.
Suggested fix
- Don't tear down/reinitialize already-connected MCP servers on a foreground-session handoff that isn't actually changing MCP config.
- If a reload is unavoidable, retry a server that was cancelled mid-handshake instead of leaving it permanently in
failedstate. - Respect the "MCP reload lock still held at disposal" case by waiting for the lock instead of draining the graph without it.
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
~/.copilot/mcp-config.json に設定された HTTP server で障害を再現し、その後、MCP reload lock 周辺のフォアグラウンドセッションのハンドオフと copilot_runtime::session::mcp::session_host のログを調査します。/clear と再起動の動作を、文書化されたキャンセルシーケンスと比較します。いずれかのハンドオフ後もリモートサーバーが接続されたままになるか、手動で再接続しなくても自動的に復旧すれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- rust
- 領域
- api, cli, networking
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 活発
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 45/100