openai / openai/codex

[Bug]: app-server stdio thread/resume emits an unbounded JSONL response that breaks bounded clients

Open
#40,362 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

app-server bug CLI
Dominant language
Rust
Stars
125k
Forks
19.4k
PR merge metrics
PR metrics pending

Description

Summary

Default thread/resume reconstructs all turns into one mandatory JSONL response. With generated images, that single stdio record can exceed a bounded client reader and make recovery impossible: the response correlation id is lost, later notifications cannot complete the pending request, and reconnect loops repeat the same oversized frame.

This is related to #21988, but the failure mode here is specifically bounded stdio framing / RPC recovery rather than only WebSocket RAM and traffic amplification. Related read_thread output suppression issue: #39148.

Linux/stdin-stdout reproduction: default thread/resume emits a 23 MB JSONL record

I reproduced the same generated-image frame amplification through the app-server's stdio transport,
where it becomes a connection/recovery failure rather than only a RAM/traffic problem.

Environment:

  • codex-cli 0.149.0
  • Linux 7.0.0-30-generic x86_64
  • codex app-server --stdio (newline-delimited JSON-RPC)
  • ChatGPT-auth production thread, replayed from a disposable clone without auth.json

The stored thread had exactly seven completed image_generation_end records. Their serialized
rollout records totaled 22,612,401 bytes; the corresponding paired large tool-output records totaled
22,608,258 bytes. No prompt, path, image bytes, or raw history was retained in the measurement.

Request sequence:

  1. initialize, request id=1
  2. initialized
  3. thread/resume, request id=2, default parameters (excludeTurns absent)

Measured stdout records:

Record Bytes
initialize response (id=1) 214
configWarning 500
remoteControl/status/changed 207
thread/status/changed 149
thread/resume response (id=2) 23,159,303

The oversized record is a response envelope (id=2, no method), not an item/turn notification.
A client with a bounded 16 MiB line reader cannot parse the envelope or resolve the pending resume
request. Discarding the record is not recovery because this is the only required response.

The existing experimental API is an effective control. With
initialize.capabilities.experimentalApi=true and thread/resume(excludeTurns=true), the exact same
thread returned a 5,104-byte id=2 response (99.978% smaller). It retained the matching thread id,
status={"type":"idle"}, preview and runtime/config metadata, and returned turns=[]. Without the
initialize capability, 0.149.0 correctly returned -32600:
thread/resume.excludeTurns requires experimentalApi capability.

This is consistent with the current source comment that thread/resume can include large MCP and
image-generation payloads, while response redaction is limited to the two ChatGPT mobile remote
client names:
https://github.com/openai/codex/blob/main/codex-rs/app-server/src/request_processors/thread_resume_redaction.rs

Expected behavior:

  • Default resume/read responses should not inline unbounded generated-image bytes into one transport
    frame.
  • Large binary results should be omitted, externalized, or referenced by saved path/content id.
  • The bounded metadata-only resume contract should remain available without reconstructing full turn
    payloads, and clients should be able to negotiate it deterministically.

Related current report: #39148 shows read_thread returning multi-megabyte imageGeneration.result
despite includeOutputs:false and maxOutputCharsPerItem.

We mitigated the client side by negotiating experimentalApi, using excludeTurns:true for stored
native resumes, and treating any lost oversized JSONL record as terminal transport corruption rather
than continuing with unresolved request correlation
(metadata-only resume and typed failure).
This avoids the crash, but it does not remove the upstream unbounded default response.

Minimal response-size scanner

Run only against a disposable CODEX_HOME containing an image-heavy test thread. It does not start a
model turn and never retains the response payload; it reports envelope id/method and byte length.

import asyncio
import json
import os
import re
import sys


async def main() -> None:
    thread_id = sys.argv[1]
    exclude_turns = "--exclude-turns" in sys.argv[2:]
    proc = await asyncio.create_subprocess_exec(
        "codex", "app-server", "--stdio",
        stdin=asyncio.subprocess.PIPE,
        stdout=asyncio.subprocess.PIPE,
        stderr=asyncio.subprocess.DEVNULL,
        env=os.environ.copy(),
    )

    async def send(payload: dict) -> None:
        proc.stdin.write((json.dumps(payload, separators=(",", ":")) + "\n").encode())
        await proc.stdin.drain()

    buffered = b""

    async def next_record() -> tuple[int | None, str | None, int]:
        nonlocal buffered
        size = 0
        prefix = bytearray()
        while True:
            if b"\n" in buffered:
                part, buffered = buffered.split(b"\n", 1)
                size += len(part) + 1
                prefix.extend(part[: max(0, 1024 - len(prefix))])
                text = prefix.decode("utf-8", "replace")
                rid = re.search(r'"id"\s*:\s*(\d+)', text)
                method = re.search(r'"method"\s*:\s*"([^"]+)"', text)
                return (
                    int(rid.group(1)) if rid else None,
                    method.group(1) if method else None,
                    size,
                )
            if buffered:
                size += len(buffered)
                prefix.extend(buffered[: max(0, 1024 - len(prefix))])
                buffered = b""
            buffered = await proc.stdout.read(65536)
            if not buffered:
                raise EOFError("app-server closed stdout")

    capabilities = {"experimentalApi": True} if exclude_turns else {}
    await send({
        "method": "initialize",
        "id": 1,
        "params": {
            "clientInfo": {"name": "resume-size-probe", "title": "probe", "version": "1"},
            "capabilities": capabilities,
        },
    })
    while (await next_record())[0] != 1:
        pass
    await send({"method": "initialized", "params": {}})
    params = {"threadId": thread_id}
    if exclude_turns:
        params["excludeTurns"] = True
    await send({"method": "thread/resume", "id": 2, "params": params})
    while True:
        record_id, method, size = await next_record()
        if record_id == 2:
            print({"id": record_id, "method": method, "bytes": size})
            break
    proc.terminate()
    await proc.wait()


asyncio.run(main())

Expected on the measured test thread:

$ CODEX_HOME=/path/to/disposable-copy python resume_size.py THREAD_ID
{'id': 2, 'method': None, 'bytes': 23159303}
$ CODEX_HOME=/path/to/disposable-copy python resume_size.py THREAD_ID --exclude-turns
{'id': 2, 'method': None, 'bytes': 5104}

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading codex-rs/app-server/src/request_processors/thread_resume_redaction.rs and the surrounding thread/resume handling, then run the provided resume-size scanner against a disposable image-heavy CODEX_HOME. Done means default resume/read responses no longer place unbounded binary results in one JSONL frame, while a deterministic metadata-only resume remains available and preserves the documented response metadata.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
api, backend, backend-api-design
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.