google / google/adk-python

MCP session stale after server scale-to-zero - @retry_on_errors reuses dead cached session

Abierto
#7,060 6 comentarios 0 reacciones 1 asignado Reclamado por @llalitkumarrr Ver en GitHub
mcp request clarification spam
Lenguaje dominante
Python
Estrellas
21.5k
Forks
4k
Merge medio
1 d 14 h
PR fusionados (30 d)
37

Descripción

## Bug Description

When an MCP server (e.g. on Cloud Run) scales to zero and back, the server's in-memory sessions are lost. The agent's `@retry_on_errors` decorator retries the failed call, but `MCPSessionManager.create_session()` returns the **same cached dead session** because `_is_session_disconnected()` only checks transport-level stream closure - the transport is still alive, only the server-side session is gone.

This means every retry reuses the stale session and fails again, until all retries are exhausted.

## Steps to Reproduce

1. Deploy an MCP server on Cloud Run (or any auto-scaling platform) with scale-to-zero enabled
2. Connect an ADK agent using `McpToolset` with a Streamable HTTP connection
3. Make a successful tool call (session is created and cached)
4. Wait for the MCP server to scale to zero (~15 minutes idle)
5. Make another tool call - the server scales back up with fresh state

## Expected Behavior

The agent should detect the server-side session loss and create a fresh session on retry.

## Actual Behavior

- `@retry_on_errors` catches the error and retries
- `create_session()` finds the cached session
- `_is_session_disconnected()` returns `False` (transport streams are still open)
- The same dead session is reused → retry fails with the same error
- All retries exhausted → permanent failure

## Root Cause

`_is_session_disconnected()` only checks `_read_stream._closed` and `_write_stream._closed`. In a scale-to-zero scenario, the HTTP transport is fine - the server responds normally with
a 404/error for the unknown session ID. The check needs a way to detect **servtion, not just transport failure.

The existing `_MCP_GRACEFUL_ERROR_HANDLING` flag handles the case where the asyncio task dies, but not this scenario where the task is alive and the server responds normally.

## Environment

- ADK version: 2.7.1
- MCP server: Cloud Run with Streamable HTTP transport
- Python: 3.12+

## Suggested Fix

Add an `invalidate_session()` method to `MCPSessionManager` that marks a session key as stale. Have `_execute_with_session()` call it on any exception before re-raising, so the next `@retry_on_errors` attempt builds a fresh session. PR incoming.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.