anthropics / anthropics/anthropic-sdk-python

[Beta Sessions / Managed Agents] MCP tool-result image content blocks reach the model as base64 text instead of vision

Đang mở
#1,671 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
3.9k
Fork
853
Merge trung bình
1 ngày 18 giờ
Pull request đã merge (30 ngày)
11

Mô tả

### Summary

In the Beta **Managed Agents / Sessions** API (self-hosted environment — Anthropic runs the agent loop; tools execute in our sandbox), when a configured **remote MCP server** returns a tool result that contains an **image content block** (`{"type":"image","data":"","mimeType":"image/jpeg"}`), the image is **not delivered to the model as vision**. The base64 payload instead reaches the model as **raw text** inside the tool result.

Two consequences:
1. **The model cannot see the image** — it can only work from any accompanying text/OCR block, and explicitly reports it "cannot see" the picture.
2. **Large token waste** — the base64 string is tokenized as text (tens of thousands of tokens per image), with no benefit.

### Expected
MCP tool-result `image` content blocks should be forwarded to the model as image (vision) input — the same way they are in a normal Claude session.

### Actual
The model receives the base64 as text.

### Isolation already done
- The **same MCP server**, called from a normal Claude session (claude.ai / Claude Code via an MCP connector), returns images that render correctly and the model can see them → the server emits valid `ImageContent` blocks.
- For Managed Agents, MCP tools execute **server-side**: the Python SDK's session runner explicitly does **not** dispatch `agent.mcp_tool_use` (in-code comment: *"MCP tools run server-side and the runner never sees a result to post for them"*). So this is in the managed-agents MCP relay, not the local tool runtime.
- `BetaManagedAgentsAgentMcpToolResultEvent`'s content union **already includes `BetaManagedAgentsImageBlock`**, so the type system supports image results — the relay appears to **flatten** `ImageContent` into a text/base64 representation instead of forwarding it as an image block to the model.

### Environment
- `anthropic` (Python) **0.109.1**
- Beta Managed Agents / Sessions, self-hosted environment

### Related
Related to but **distinct from** #1329: that issue is about the **claude.ai UI** not rendering tool images **inline to the user** (where *"Claude can see the image"*). This report is the opposite — in Managed Agents the **model itself does not receive the image** as vision.

### Repro
Configure a managed agent with a remote MCP server whose tool returns `[TextContent, ImageContent]`; run a session that calls that tool; observe that the model receives the base64 as text rather than seeing the image. Happy to provide a minimal MCP server and a session id on request.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.