modelcontextprotocol / modelcontextprotocol/java-sdk

SSE client silently truncates a data: line at U+2028/U+2029/U+0085, then fails with "Error parsing JSON-RPC message"

Abierto Apto para principiantes
#1,136 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

area/transport bug P2
Lenguaje dominante
Java
Estrellas
3.7k
Forks
1.1k
Merge medio
1 d 15 h
PR fusionados (30 d)
9

Descripción

Summary

ResponseSubscribers.SseLineSubscriber extracts the payload of an SSE data: line with a MULTILINE regex. Java's MULTILINE mode treats (LINE SEPARATOR), (PARAGRAPH SEPARATOR) and … (NEL) as line terminators, so when a data: line contains one of those characters the capture group stops there and everything after it is silently discarded. The client then fails to deserialise the truncated JSON and throws:

io.modelcontextprotocol.spec.McpTransportException: Error parsing JSON-RPC message: SseResponseEvent[...]

Any tool result, resource content or prompt text containing one of these three characters is unreadable by the client. They are legal unescaped inside a JSON string, and they turn up in real content — text pasted from word processors, web pages and PDFs.

Affected code

mcp-core/src/main/java/io/modelcontextprotocol/client/transport/ResponseSubscribers.java:

private static final Pattern EVENT_DATA_PATTERN = Pattern.compile("^data:(.+)$", Pattern.MULTILINE);
private static final Pattern EVENT_ID_PATTERN   = Pattern.compile("^id:(.+)$",   Pattern.MULTILINE);
private static final Pattern EVENT_TYPE_PATTERN = Pattern.compile("^event:(.+)$", Pattern.MULTILINE);
if (line.startsWith("data:")) {
    var matcher = EVENT_DATA_PATTERN.matcher(line);
    if (matcher.find()) {
        String data = matcher.group(1).trim();
        ...
        this.eventBuilder.append(data).append("\n");
    }
    upstream().request(1);
}

Two regex properties combine here:

  • . without DOTALL does not match \n, \r, …, or .
  • $ in MULTILINE mode matches before any of those.

So ^data:(.+)$ happily matches a prefix of the line and find() returns true — the truncation is not detectable at the match site.

The line splitter feeding the subscriber (HttpResponse.BodySubscribers.fromLineSubscriber, which uses BufferedReader.readLine() semantics) splits on \n, \r and \r\n only. //… therefore arrive inside a line, where only the regex sees them.

This affects both client transports that use the subscriber: HttpClientStreamableHttpTransport and HttpClientSseClientTransport.

Reproducer

Isolating the regex (JDK 25, but the behaviour is not version-specific):

import java.util.regex.*;

public class SseProbe {
    static final Pattern P = Pattern.compile("^data:(.+)$", Pattern.MULTILINE);

    static void probe(String name, char c) {
        String line = "data:{\"text\":\"" + c + "tail\"}";
        Matcher m = P.matcher(line);
        String got = m.find() ? m.group(1) : "<no match>";
        System.out.printf("%-8s U+%04X  line.len=%d  captured=%s%n",
                name, (int) c, line.length(), got.replace(String.valueOf(c), "<CHAR>"));
    }

    public static void main(String[] args) {
        probe("LS", '
');
        probe("PS", '
');
        probe("NEL", '…');
        probe("VT", '');
        probe("plain", 'X');
    }
}
LS       U+2028  line.len=21  captured={"text":"
PS       U+2029  line.len=21  captured={"text":"
NEL      U+0085  line.len=21  captured={"text":"
VT       U+000B  line.len=21  captured={"text":"<CHAR>tail"}
plain    U+0058  line.len=21  captured={"text":"<CHAR>tail"}

End to end: have a server return a CallToolResult whose text content contains "a
b", and call it over HttpClientStreamableHttpTransport. The client throws McpTransportException: Error parsing JSON-RPC message instead of returning the result.

Impact in production

We run a gateway that proxies several remote MCP servers. Over the 30 days to 2026-09-18 we logged 179 failed tool calls with this exception, all against one backend (the hosted Atlassian MCP server) — Confluence comment bodies and Jira issue descriptions that contain one of these characters. From the caller's side the tool simply looks broken, and it is reproducible per document: the same tool succeeds on every other page.

Suggested fix

Drop the regexes and strip the field prefix directly, which is also what the SSE spec describes — collect the characters after the colon, removing a single leading U+0020 if present:

if (line.startsWith("data:")) {
    String data = line.substring(5);
    if (data.startsWith(" ")) {
        data = data.substring(1);
    }
    ...
}

Note that the current .trim() on the captured group is lossier than the spec allows in any case: it strips all leading and trailing whitespace rather than the single optional space, so a payload with significant leading/trailing whitespace is also altered.

A regression test wants a data: line containing and an assertion that the emitted SseEvent.data() is byte-identical to what was written.

Environment

  • io.modelcontextprotocol.sdk:mcp 1.1.0; also present unchanged on main (2.1.0-SNAPSHOT, checked 2026-09-18)
  • JDK 25

Possibly of interest to whoever picks this up: #1042 is a separate problem in the same subscriber.

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Comienza en mcp-core/src/main/java/io/modelcontextprotocol/client/transport/ResponseSubscribers.java y sigue cómo SseLineSubscriber maneja las líneas data: para ambos transportes afectados. Añade la prueba de regresión descrita en el issue y verifica después que SseEvent.data() conserva el contenido U+2028, U+2029 y U+0085, y que no se eliminan los espacios en blanco circundantes significativos.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
java
Área
networking
Tipo de issue
Error
Dificultad
2/5
Tiempo estimado
1-3 horas
Estado de actividad
Activo
Claridad
Bien especificado
Aptitud para principiantes
88/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.