modelcontextprotocol / modelcontextprotocol/java-sdk

StdioClientTransport missing explicit UTF-8 charset in InputStreamReader (same issue as #295, but on client side)

未關閉 適合新手
#898 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

bug P2 ready for work
主要語言
Java
星號
3.7k
分支
1.1k
平均合併
1 天 15 小時
30 天內合併 PR
9

描述

Bug description

StdioClientTransport has the same encoding mismatch issue that was identified in #295 and fixed for StdioServerTransportProvider in #826 — but the fix was only applied to the server side. The client transport still lacks explicit UTF-8 charset
specification when reading from the subprocess.

In startInboundProcessing, the InputStreamReader is created without specifying a charset:

https://github.com/modelcontextprotocol/java-sdk/blob/main/mcp-core/src/main/java/io/modelcontextprotocol/client/transport/StdioClientTransport.java#L249

try (BufferedReader processReader = new BufferedReader(new InputStreamReader(process.getInputStream()))) {

Similarly, in startErrorProcessing:

https://github.com/modelcontextprotocol/java-sdk/blob/main/mcp-core/src/main/java/io/modelcontextprotocol/client/transport/StdioClientTransport.java#L182-L183

try (BufferedReader processErrorReader = new BufferedReader(
        new InputStreamReader(process.getErrorStream()))) {

Meanwhile, startOutboundProcessing already correctly specifies UTF-8:

os.write(jsonMessage.getBytes(StandardCharsets.UTF_8));
os.write("\n".getBytes(StandardCharsets.UTF_8));

This is the exact same inconsistency that #295 reported for StdioServerTransportProvider, and that #826 fixed — only on the server side.

Steps to reproduce

  1. Start a JVM with default charset set to something other than UTF-8 (e.g., -Dfile.encoding=COMPAT on Windows with Japanese locale, which resolves to MS932/Shift_JIS)
  2. Connect to an MCP server via StdioClientTransport
  3. Call a tool that returns multi-byte UTF-8 characters (e.g., Japanese, Chinese, Korean, emoji) in its response

Expected behavior

Multi-byte characters in the server's JSON-RPC response should be decoded correctly, since the MCP stdio transport specification requires UTF-8.

Actual behavior

The InputStreamReader uses Charset.defaultCharset() instead of UTF-8. When the default charset is not UTF-8, the response bytes are decoded with the wrong charset, corrupting multi-byte characters. This corruption can also break the JSON structure itself,
resulting in JsonParseException:

com.fasterxml.jackson.core.JsonParseException: Unexpected character ('' (code 92)): was expecting double-quote to start field name

For example, with MS932 as the default charset, the last byte of certain UTF-8 characters (0x8B, etc.) is interpreted as a MS932 lead byte, which then consumes the following byte — potentially a JSON structural character like \ (0x5C). This shifts the parser
state and breaks JSON parsing entirely.

Environment

  • MCP Java SDK version: 1.1.1
  • Java version: 21
  • OS: Windows 11 (Japanese locale, default charset MS932 with -Dfile.encoding=COMPAT)

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 mcp-core/src/main/java/io/modelcontextprotocol/client/transport/StdioClientTransport.java 開始,閱讀 startErrorProcessing 和 startInboundProcessing,並參考 startOutboundProcessing 中現有的 UTF-8 處理。比較 #826 中相關的伺服器端修正,並確認當 JVM 預設字元集不是 UTF-8 時,回應和錯誤仍能正確解碼。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
java
領域
backend
Issue 類型
缺陷
難度
2/5
預估耗時
1-3 小時
活躍度
冷清
描述清晰度
描述清楚
新手友好度
74/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。