Remote/HTTP MCP servers (e.g. atlassian) are stranded 'failed' after every /clear or session relaunch
Nadie ha tomado este issue todavía.
- Lenguaje dominante
- Shell
- Estrellas
- 11.2k
- Forks
- 1.9k
- Merge medio
- 14 h 16 min
- PR fusionados (30 d)
- 6
Descripción
Summary
Remote (HTTP) MCP server connections are reliably stranded in a failed state whenever the CLI performs a foreground-session handoff — which happens on every /clear, and apparently on session resume/relaunch too. The MCP connection graph is torn down and rebuilt during the handoff, and slower remote/OAuth-backed servers lose the reconnect race and never recover automatically.
Environment
- CLI version: 1.0.83 (macOS, confirmed already latest via in-app update check)
- MCP server affected:
atlassian(https://mcp.atlassian.com/v2/mcp, HTTP transport with Authorization header) - Other MCP servers configured (
github-mcp-server,codegraph,local-rag— local/stdio) are not visibly affected, likely because their reconnect is fast enough to win the race.
Steps to reproduce
- Start a Copilot CLI session with the
atlassianMCP server configured (~/.copilot/mcp-config.json, HTTP transport). - Confirm it connects fine (
/mcp show atlassian→ connected, tools listed). - Run
/clear(or relaunch/resume a session). - Run
/mcp show atlassianagain.
Expected
atlassian reconnects cleanly like the other MCP servers.
Actual
atlassian is left in:
{
"name": "atlassian",
"status": "failed",
"error": "MCP server \"atlassian\" connection was cancelled"
}
It does not self-heal — it requires a manual reconnect action or a full CLI restart, and even a full restart reproduces the same failure deterministically (confirmed across two separate relaunches, ~20 minutes apart).
Root cause (from ~/.copilot/logs/process-*.log)
Every /clear/relaunch briefly registers a throwaway foreground session, then immediately unregisters it in favor of the real one, within milliseconds:
17:27:22.425Z [INFO] Registering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.717Z [INFO] Unregistering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.720Z [INFO] Registering foreground session: d89d1d9a-99c4-4247-86fe-a75b4da5daad
17:27:28.783Z [INFO] Closing session 6266dbfb-39da-4797-9265-c37d9655a3fd
This handoff triggers a full MCP graph reload while a reload lock is still held from the previous teardown:
17:05:05.167Z [WARNING] [rust:copilot_runtime::session::mcp::session_host] MCP reload lock still held at disposal; draining the graph without it
Every MCP server — including atlassian — gets re-initialized and then almost immediately cancelled mid-handshake:
17:27:26.104Z [INFO] [rust:rmcp::service] Service initialized as client {... "name": "atlassian-mcp-server" ...}
17:27:28.809Z [INFO] [rust:rmcp::service] task cancelled
17:27:28.809Z [INFO] [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
Local/stdio servers appear to reconnect fast enough afterward to recover; atlassian's remote HTTP+OAuth handshake is slower and consistently loses the race, leaving it permanently failed with no automatic retry.
Confirmed not a network/auth issue independently:
- DNS resolves fine for
mcp.atlassian.com. curl https://mcp.atlassian.com/v2/mcpreturns401(server reachable, no valid header supplied in the manual test — expected).ATLASSIAN_MCP_AUTH_HEADERenv var is set.
Impact
Every /clear (a very common action) breaks the Atlassian MCP integration and requires a manual reconnect or CLI restart to restore it — significant daily friction for anyone using an HTTP-based MCP server.
Suggested fix
- Don't tear down/reinitialize already-connected MCP servers on a foreground-session handoff that isn't actually changing MCP config.
- If a reload is unavoidable, retry a server that was cancelled mid-handshake instead of leaving it permanently in
failedstate. - Respect the "MCP reload lock still held at disposal" case by waiting for the lock instead of draining the graph without it.
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Línea de trabajo
Reproduce el fallo con un HTTP server configurado en ~/.copilot/mcp-config.json y, a continuación, inspecciona los logs de la transferencia de la sesión en primer plano y de copilot_runtime::session::mcp::session_host alrededor del MCP reload lock. Compara el comportamiento de /clear y del relanzamiento con la secuencia de cancelación documentada. Se considera terminado cuando el servidor remoto permanece conectado o se recupera automáticamente después de cualquiera de las dos transferencias, sin requerir una reconexión manual.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- rust
- Área
- api, cli, networking
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Activo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 45/100