MCP tool result image content blocks are not surfaced to the model
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Shell
- Star
- 11.2k
- Fork
- 1.9k
- Merge trung bình
- 14 giờ 16 phút
- Pull request đã merge (30 ngày)
- 6
Mô tả
Describe the bug
When an MCP server returns an image content block in a tool result (per the MCP specification: content array item {type:"image", mimeType:"image/png", data:}), the image never reaches the model. Only the text blocks and structuredContent of the same tool result are delivered.
Verified from the agent side (Copilot CLI 1.0.80, Windows):
- The MCP server's wire format is spec-compliant. Capturing the raw stdio JSON-RPC response for a screenshot tool shows content[0] = text block and content[1] = a valid image block (mimeType image/png, correct base64 PNG payload).
-
- Across multiple calls (window screenshot, region screenshot), the image data never appeared in the model context; only the text metadata block arrived.
-
- Static inspection of app.js (1.0.80) shows image handling only in two paths: user-attachment sanitization (file_data / image_url / input_image) and the mcp-sampling converter for user-role messages. No conversion path exists for MCP tool-result image blocks, and tool results are assembled through the native session layer (session.mcp.apps.callTool).
User-attached images (paste / drag-drop / @file) work fine, so the model and plan support vision. The gap is specific to images returned by MCP tools.
- Static inspection of app.js (1.0.80) shows image handling only in two paths: user-attachment sanitization (file_data / image_url / input_image) and the mcp-sampling converter for user-role messages. No conversion path exists for MCP tool-result image blocks, and tool results are assembled through the native session layer (session.mcp.apps.callTool).
Affected version
1.0.80 (Windows 10.0.26200, PowerShell)
Steps to reproduce the behavior
- Configure any MCP server whose tool returns an image content block (e.g. a UI-automation server with a screenshot tool; any tool returning {type:"image", mimeType:"image/png", data:} reproduces it).
-
- In an interactive Copilot CLI session, ask the agent to call that tool and then describe what it sees in the image.
-
- Observe: the agent only receives the text blocks / structuredContent of the tool result. It reports the image metadata but cannot see the actual pixels, and often states the image content is unavailable.
Expected behavior
When the selected model supports vision, image content blocks inside MCP tool results should be passed to the model (like user-attached images), so agents can act on screenshots and other tool-generated images.
Additional context
This blocks a whole class of computer-use / UI-automation workflows: MCP servers that follow the spec and return screenshots as image blocks work in other agent CLIs that surface tool-result images, but not in Copilot CLI. Current workaround: agents must rely on structured reads only, or the user must manually attach the image.
Issue submitted by GLM-5.3 via rpacu MCP
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Bắt đầu bằng cách truy vết app.js từ phần xử lý hình ảnh hiện có của tệp đính kèm người dùng và MCP-sampling đến các kết quả công cụ được tập hợp thông qua session.mcp.apps.callTool. Tái hiện bằng một công cụ chụp màn hình MCP trả về một khối hình ảnh tuân thủ đặc tả, sau đó xác minh rằng một mô hình có khả năng xử lý hình ảnh nhận được hình ảnh và có thể mô tả hình ảnh đó, thay vì chỉ nhận siêu dữ liệu dạng văn bản của nó.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- javascript, shell
- Lĩnh vực
- ai, cli
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Sôi nổi
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 62/100