google / google/adk-python

MCP session stale after server scale-to-zero - @retry_on_errors reuses dead cached session

Ouverte
#7,060 6 commentaires 0 réactions 1 personne assignée Réclamée par @llalitkumarrr Voir sur GitHub
mcp request clarification spam
Langage dominant
Python
Étoiles
21.5k
Forks
4k
Merge moyen
1 j 14 h
PR mergées (30 j)
37

Description

## Bug Description

When an MCP server (e.g. on Cloud Run) scales to zero and back, the server's in-memory sessions are lost. The agent's `@retry_on_errors` decorator retries the failed call, but `MCPSessionManager.create_session()` returns the **same cached dead session** because `_is_session_disconnected()` only checks transport-level stream closure - the transport is still alive, only the server-side session is gone.

This means every retry reuses the stale session and fails again, until all retries are exhausted.

## Steps to Reproduce

1. Deploy an MCP server on Cloud Run (or any auto-scaling platform) with scale-to-zero enabled
2. Connect an ADK agent using `McpToolset` with a Streamable HTTP connection
3. Make a successful tool call (session is created and cached)
4. Wait for the MCP server to scale to zero (~15 minutes idle)
5. Make another tool call - the server scales back up with fresh state

## Expected Behavior

The agent should detect the server-side session loss and create a fresh session on retry.

## Actual Behavior

- `@retry_on_errors` catches the error and retries
- `create_session()` finds the cached session
- `_is_session_disconnected()` returns `False` (transport streams are still open)
- The same dead session is reused → retry fails with the same error
- All retries exhausted → permanent failure

## Root Cause

`_is_session_disconnected()` only checks `_read_stream._closed` and `_write_stream._closed`. In a scale-to-zero scenario, the HTTP transport is fine - the server responds normally with
a 404/error for the unknown session ID. The check needs a way to detect **servtion, not just transport failure.

The existing `_MCP_GRACEFUL_ERROR_HANDLING` flag handles the case where the asyncio task dies, but not this scenario where the task is alive and the server responds normally.

## Environment

- ADK version: 2.7.1
- MCP server: Cloud Run with Streamable HTTP transport
- Python: 3.12+

## Suggested Fix

Add an `invalidate_session()` method to `MCPSessionManager` that marks a session key as stale. Have `_execute_with_session()` call it on any exception before re-raising, so the next `@retry_on_errors` attempt builds a fresh session. PR incoming.

Guide de contribution

Ouvrir le guide de contribution

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.