modelcontextprotocol / modelcontextprotocol/java-sdk

SSE client silently truncates a data: line at U+2028/U+2029/U+0085, then fails with "Error parsing JSON-RPC message"

Ouverte Adaptée aux débutants
#1,136 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

area/transport bug P2
Langage dominant
Java
Étoiles
3.7k
Forks
1.1k
Merge moyen
1 j 15 h
PR mergées (30 j)
9

Description

Summary

ResponseSubscribers.SseLineSubscriber extracts the payload of an SSE data: line with a MULTILINE regex. Java's MULTILINE mode treats (LINE SEPARATOR), (PARAGRAPH SEPARATOR) and … (NEL) as line terminators, so when a data: line contains one of those characters the capture group stops there and everything after it is silently discarded. The client then fails to deserialise the truncated JSON and throws:

io.modelcontextprotocol.spec.McpTransportException: Error parsing JSON-RPC message: SseResponseEvent[...]

Any tool result, resource content or prompt text containing one of these three characters is unreadable by the client. They are legal unescaped inside a JSON string, and they turn up in real content — text pasted from word processors, web pages and PDFs.

Affected code

mcp-core/src/main/java/io/modelcontextprotocol/client/transport/ResponseSubscribers.java:

private static final Pattern EVENT_DATA_PATTERN = Pattern.compile("^data:(.+)$", Pattern.MULTILINE);
private static final Pattern EVENT_ID_PATTERN   = Pattern.compile("^id:(.+)$",   Pattern.MULTILINE);
private static final Pattern EVENT_TYPE_PATTERN = Pattern.compile("^event:(.+)$", Pattern.MULTILINE);
if (line.startsWith("data:")) {
    var matcher = EVENT_DATA_PATTERN.matcher(line);
    if (matcher.find()) {
        String data = matcher.group(1).trim();
        ...
        this.eventBuilder.append(data).append("\n");
    }
    upstream().request(1);
}

Two regex properties combine here:

  • . without DOTALL does not match \n, \r, …, or .
  • $ in MULTILINE mode matches before any of those.

So ^data:(.+)$ happily matches a prefix of the line and find() returns true — the truncation is not detectable at the match site.

The line splitter feeding the subscriber (HttpResponse.BodySubscribers.fromLineSubscriber, which uses BufferedReader.readLine() semantics) splits on \n, \r and \r\n only. //… therefore arrive inside a line, where only the regex sees them.

This affects both client transports that use the subscriber: HttpClientStreamableHttpTransport and HttpClientSseClientTransport.

Reproducer

Isolating the regex (JDK 25, but the behaviour is not version-specific):

import java.util.regex.*;

public class SseProbe {
    static final Pattern P = Pattern.compile("^data:(.+)$", Pattern.MULTILINE);

    static void probe(String name, char c) {
        String line = "data:{\"text\":\"" + c + "tail\"}";
        Matcher m = P.matcher(line);
        String got = m.find() ? m.group(1) : "<no match>";
        System.out.printf("%-8s U+%04X  line.len=%d  captured=%s%n",
                name, (int) c, line.length(), got.replace(String.valueOf(c), "<CHAR>"));
    }

    public static void main(String[] args) {
        probe("LS", '
');
        probe("PS", '
');
        probe("NEL", '…');
        probe("VT", '');
        probe("plain", 'X');
    }
}
LS       U+2028  line.len=21  captured={"text":"
PS       U+2029  line.len=21  captured={"text":"
NEL      U+0085  line.len=21  captured={"text":"
VT       U+000B  line.len=21  captured={"text":"<CHAR>tail"}
plain    U+0058  line.len=21  captured={"text":"<CHAR>tail"}

End to end: have a server return a CallToolResult whose text content contains "a
b", and call it over HttpClientStreamableHttpTransport. The client throws McpTransportException: Error parsing JSON-RPC message instead of returning the result.

Impact in production

We run a gateway that proxies several remote MCP servers. Over the 30 days to 2026-09-18 we logged 179 failed tool calls with this exception, all against one backend (the hosted Atlassian MCP server) — Confluence comment bodies and Jira issue descriptions that contain one of these characters. From the caller's side the tool simply looks broken, and it is reproducible per document: the same tool succeeds on every other page.

Suggested fix

Drop the regexes and strip the field prefix directly, which is also what the SSE spec describes — collect the characters after the colon, removing a single leading U+0020 if present:

if (line.startsWith("data:")) {
    String data = line.substring(5);
    if (data.startsWith(" ")) {
        data = data.substring(1);
    }
    ...
}

Note that the current .trim() on the captured group is lossier than the spec allows in any case: it strips all leading and trailing whitespace rather than the single optional space, so a payload with significant leading/trailing whitespace is also altered.

A regression test wants a data: line containing and an assertion that the emitted SseEvent.data() is byte-identical to what was written.

Environment

  • io.modelcontextprotocol.sdk:mcp 1.1.0; also present unchanged on main (2.1.0-SNAPSHOT, checked 2026-09-18)
  • JDK 25

Possibly of interest to whoever picks this up: #1042 is a separate problem in the same subscriber.

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez dans mcp-core/src/main/java/io/modelcontextprotocol/client/transport/ResponseSubscribers.java et suivez la gestion des lignes data: par SseLineSubscriber pour les deux transports concernés. Ajoutez le test de régression décrit dans l’issue, puis vérifiez que SseEvent.data() préserve le contenu U+2028, U+2029 et U+0085 et que les espaces blancs environnants significatifs ne sont pas supprimés.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
java
Domaine
networking
Type d'issue
Bug
Difficulté
2/5
Temps estimé
1-3 heures
Activité
Active
Clarté
Clairement spécifiée
Accessibilité débutants
88/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.