openai / openai/codex

Feature request: deliver code-mode `notify()` updates as messages the model can attribute

Open
#41,820 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

CLI enhancement tool-calls
Dominant language
Rust
Stars
125k
Forks
19.5k
PR merge metrics
PR metrics pending

Description

What variant of Codex are you using?

CLI

What feature would you like to see?

TL;DR

Code-mode's JS notify() injects each late update as an extra custom_tool_call_output linked to its exec call only by call_id — a field that never reaches the model's token stream, so attribution degrades into arrival-order guessing. In a controlled request with two parallel exec cells, the model swapped the two tasks' final results (foo/bar answered as 44/42 instead of 42/44) while getting every text-carried association right. We propose delivering out-of-band notify() updates as standalone user (or developer) message items whose text carries the cell/call marker — putting the association in content, the only channel that survives rendering.

You can run last two examples, Bug Demo and Expected Demo to see what happens.

Summary

Code-mode's JS notify() helper currently injects each update as an extra custom_tool_call_output hanging on the original exec call's call_id. We propose changing the delivery of these out-of-band updates so that their association to the originating call/cell travels in message content — concretely, as standalone user (or developer) message items carrying an explicit cell/call marker.

Today the association exists only in the call_id structure, and call_id never reaches the model's token stream. With two parallel exec cells, the model demonstrably misattributes the updates — it swapped the final results of the two tasks in a controlled request.

Background

Code-mode exec is the only codex tool that can produce multiple outputs for a single tool call:

  1. The model calls exec; after yield_time_ms the call yields with Script running with cell ID {cell_id} (first output, 1:1 with the call).
  2. While the cell keeps running, the script can call the JS notify(value) helper. Each call injects an extra custom_tool_call_output item with the same call_id into the conversation (codex-rs/core/src/tools/code_mode/delegate.rs:373):

https://github.com/openai/codex/blob/d58d0e5841e0de08e251673db2d5af8cf3a1ad51/codex-rs/core/src/tools/code_mode/delegate.rs#L373-L390

Every other tool path is strictly 1:1 (wait is a new call each time; yield is the current call's own response; aborted outputs are single repair placeholders). So the multi-output shape exists only through notify.

Codex already follows the "association lives in content" rule elsewhere: the yield output spells out Script running with cell ID {cell_id} (codex-rs/core/src/tools/code_mode/mod.rs:303), and wait takes cell_id as a readable argument. notify is the one path that doesn't — the text is injected verbatim, with no cell/call marker.

Evidence: attribution fails between parallel cells

The attached request body replays a self-consistent history against the Responses API. The exec tool definition in it documents the real code-mode API surface (tools.exec_command(...), the notify() helper, yield/cell-ID semantics), and the two scripts produce every output item in the replayed history — each custom_tool_call_output is traceable to a concrete notify() call in the script source.

The scenario: the user asks to run both projects' test suites in parallel with progress updates. The model issues two exec calls with this script (call_1 shown; call_2 is identical with cwd: "/workspace/bar"):

// @exec: {"yield_time_ms": 1000}
const t0 = Date.now();
const list = await tools.exec_command({cmd: "ls tests/test_*.py", cwd: "/workspace/foo"});
const files = list.trim().split("\n");
let passed = 0;
for (const f of files) {
  const out = await tools.exec_command({cmd: `pytest -q ${f}`, cwd: "/workspace/foo", yield_time_ms: 60000});
  passed += Number(out.match(/(\d+) passed/)?.[1] ?? 0);
  notify(`Script Running. Finish ${passed} tasks`);
}
const finalOut = await tools.exec_command({cmd: "pytest -q tests/", cwd: "/workspace/foo", yield_time_ms: 120000});
const total = Number(finalOut.match(/(\d+) passed/)?.[1] ?? passed);
const secs = ((Date.now() - t0) / 1000).toFixed(1);
notify(`Script completed\nWall time ${secs} seconds\nOutput:\n${total} passed`);

Both scripts yield after 1s (Script running with cell ID cell-42 / cell-43), then keep running. /workspace/foo has 6 test files → 6 progress notifications from call_1 (cumulative counts 1, 3, 13, 22, 24, 31); /workspace/bar has 4 → 4 from call_2 (2, 23, 30, 41). The progress totals stay below the final totals because each script ends with a full-suite run whose count is the authoritative one. All of these outputs land in the conversation interleaved, mixed with an unrelated user/assistant exchange.

Final completion outputs, in arrival order: call_2 first ("44 passed"), then call_1 ("42 passed").

Trailing user question: "Have the backend tests finished running? What were the results? How many tests were successfully completed for each project?"

Ground truth:

call project final result
call_1 (cell-42) /workspace/foo 42 passed
call_2 (cell-43) /workspace/bar 44 passed
Observed behavior

A run of this scenario (store: false, stream: true), using an earlier presentation of the identical history (the attached body is the refined, fully self-consistent version — same call_ids, output sequence, ordering, and question, with two cosmetic text differences: the completion texts were trimmed from N passed in 186.9s to N passed so the script templates reproduce them exactly, and call_2's last progress count was changed from 42 to 41 to avoid colliding with call_1's 42 passed):

model's answer for foo model's answer for bar
44 passed ❌ 42 passed ❌

The run got the numbers right but swapped the projects.

Telling detail: the model attributed the cell IDs correctly (cell-42 ↔ foo, cell-43 ↔ bar) — because that association is spelled out in the output text right next to the calls. It only swapped the final results, whose association exists solely via call_id. The error pattern is consistent with FIFO guessing: the completion that arrived first ("44 passed") was assigned to the call that started first (call_1).

Why the association has to live in content (format level)

The model never sees the JSON items as written — they are rendered into a token stream before reaching it. The harmony format (the public response format used by gpt-oss) shows what survives that rendering:

  1. call_id does not exist in the message model. harmony's Message struct has exactly five fields — author, recipient, content, channel, content_type (src/chat.rs:106-138). There is no field where a call_id could land; the concept simply does not exist at this layer.

  2. Tool outputs render with only a tool name and a position. The template is:

<|start|>{toolname} to=assistant<|channel|>commentary<|message|>{output}<|end|>

An output is linked to its call by (a) the tool name in the header and (b) appearing after the corresponding <|call|> message. Nothing else.

  1. Parallel calls to the same tool therefore render as indistinguishable messages. The replayed history becomes:
<|start|>assistant to=functions.exec<|channel|>commentary<|message|>{script_1}<|call|>
<|start|>assistant to=functions.exec<|channel|>commentary<|message|>{script_2}<|call|>
<|start|>functions.exec to=assistant<|message|>{out_call_1_progress}<|end|>
<|start|>functions.exec to=assistant<|message|>{out_call_2_progress}<|end|>
<|start|>functions.exec to=assistant<|message|>{out_call_1_completion}<|end|>

Every output message is structurally identical; only content and position differ. Whether out_call_1_completion belongs to the first or the second call is unanswerable from the token stream, so arrival-order (FIFO) guessing is the only attribution strategy available — which is precisely the error the model made.

  1. Nothing validates pairing. harmony's renderer walks the message list one by one with no call/output pairing checks; the documented format only ever describes one tool message per call. "Multiple outputs for one call" is a shape this format has no representation for.

Caveat: harmony is the public format of gpt-oss; the serving format of GPT-5-class models is not public. But the observed behavior matches harmony's structural predictions point for point — text-carried associations (cell IDs) are used correctly, structure-carried ones (call_id) are not, and the failure mode is arrival-order guessing.

Consequence: any fix must act above the rendering layer. The only channel that survives rendering is message content — which is exactly what the proposal below puts the association into.

Proposal

Deliver out-of-band notify() updates as user (or developer) message items instead of tool outputs. The further an output drifts from its call, the less the call_id structure seems to link it back — so late notify() updates would become standalone message items whose text carries the association, e.g.

{"type": "message", "role": "user",
 "content": [{"type": "input_text",
              "text": "<notify cell_id=\"cell-42\">Script Running. Finish 31 tasks</notify>"}]}

The call's first (yield) output stays a normal 1:1 tool output; only the late updates change shape. This keeps history append-only (no edits), removes the second-output-per-call shape entirely, and mirrors how other production agents surface background-task notifications to the model.

Alternatives considered

Envelope the injected text in place (delegate.rs:379), e.g. [cell cell-42] <text> — a one-line change that also puts the association into content, and would fix the symptom today. We still prefer the proposal above, because the envelope option preserves the fiction: the output keeps hanging on a call_id, a shape that quietly tells every reader — and every future maintainer — that the call_id linkage does something. It doesn't: call_id never reaches the model at all, so any reliance on it is reliance on a channel that isn't there. The proposal removes the fiction along with the symptom: with exactly one tool output per call, there is no structural channel left to mistakenly trust — the association exists only where the model can actually see it, in content.

Expected outcome
  • In the replayed scenario above, the model answers with the correct per-project results (foo: 42, bar: 44) because each update names its cell in text — demonstrated in Expected Demo below.
  • Every exec call keeps exactly one tool output (its 1:1 yield/result); attribution of late updates no longer depends on call_id or position.
  • No change to any 1:1 tool path (exec yield/result, wait, abort repair).
Environment
Bug Demo
{
  "model": "gpt-5.6-sol",
  "input": [
    {
      "type": "additional_tools",
      "role": "developer",
      "tools": [
        {
          "type": "custom",
          "name": "exec",
          "description": "Run JavaScript code to orchestrate/compose tool calls\n- Evaluates the provided JavaScript code in a fresh V8 isolate as an async module.\n- All nested tools are available on the global `tools` object, for example `await tools.exec_command(...)`.\n- Nested tool methods take either a string or an object as their input argument.\n- Nested tools return either an object or a string, based on the description.\n- Runs raw JavaScript -- no Node, no file system, no network access, no console.\n- You may optionally start the tool input with a first-line pragma like `// @exec: {\"yield_time_ms\": 10000, \"max_output_tokens\": 1000}`.\n- `yield_time_ms` asks `exec` to yield early if the script is still running. Defaults to 10000 ms. When it yields, the result says `Script running with cell ID ...` and the script keeps running in the background.\n- Global helpers:\n- `text(value: any)`: Appends a text item to the output.\n- `notify(value: any)`: immediately injects an extra `custom_tool_call_output` for the current `exec` call. Use it to report progress from a long-running script.\n- `store(key: string, value: any)` / `load(key: string)`: persist values between `exec` calls in the same session.\n- `setTimeout(callback: () => void, delayMs?: number)`: schedules a callback; pending timeouts do not keep `exec` alive by themselves.\n- `yield_control()`: yields the accumulated output to the model immediately while the script keeps running.\n\nNested tools:\n- `exec_command(args: {cmd: string, cwd?: string, yield_time_ms?: number, max_output_tokens?: number}): Promise<string>` -- Execute a shell command and return its output.",
          "format": {
            "type": "grammar",
            "syntax": "lark",
            "definition": "start: /(.|\\n)*/"
          }
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Run the test suites of both /workspace/foo and /workspace/bar for me at the same time, and keep me posted on the progress."
        }
      ]
    },
    {
      "type": "custom_tool_call",
      "call_id": "call_1",
      "name": "exec",
      "input": "// @exec: {\"yield_time_ms\": 1000}\nconst t0 = Date.now();\nconst list = await tools.exec_command({cmd: \"ls tests/test_*.py\", cwd: \"/workspace/foo\"});\nconst files = list.trim().split(\"\\n\");\nlet passed = 0;\nfor (const f of files) {\n  const out = await tools.exec_command({cmd: `pytest -q ${f}`, cwd: \"/workspace/foo\", yield_time_ms: 60000});\n  passed += Number(out.match(/(\\d+) passed/)?.[1] ?? 0);\n  notify(`Script Running. Finish ${passed} tasks`);\n}\nconst finalOut = await tools.exec_command({cmd: \"pytest -q tests/\", cwd: \"/workspace/foo\", yield_time_ms: 120000});\nconst total = Number(finalOut.match(/(\\d+) passed/)?.[1] ?? passed);\nconst secs = ((Date.now() - t0) / 1000).toFixed(1);\nnotify(`Script completed\\nWall time ${secs} seconds\\nOutput:\\n${total} passed`);"
    },
    {
      "type": "custom_tool_call",
      "call_id": "call_2",
      "name": "exec",
      "input": "// @exec: {\"yield_time_ms\": 1000}\nconst t0 = Date.now();\nconst list = await tools.exec_command({cmd: \"ls tests/test_*.py\", cwd: \"/workspace/bar\"});\nconst files = list.trim().split(\"\\n\");\nlet passed = 0;\nfor (const f of files) {\n  const out = await tools.exec_command({cmd: `pytest -q ${f}`, cwd: \"/workspace/bar\", yield_time_ms: 60000});\n  passed += Number(out.match(/(\\d+) passed/)?.[1] ?? 0);\n  notify(`Script Running. Finish ${passed} tasks`);\n}\nconst finalOut = await tools.exec_command({cmd: \"pytest -q tests/\", cwd: \"/workspace/bar\", yield_time_ms: 120000});\nconst total = Number(finalOut.match(/(\\d+) passed/)?.[1] ?? passed);\nconst secs = ((Date.now() - t0) / 1000).toFixed(1);\nnotify(`Script completed\\nWall time ${secs} seconds\\nOutput:\\n${total} passed`);"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script running with cell ID cell-42\nWall time 1.0 seconds\nOutput:\n(collecting...)"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_2",
      "output": "Script running with cell ID cell-43\nWall time 1.0 seconds\nOutput:\n(collecting...)"
    },
    {
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "Both test suites are now running in the background (cell-42 for /workspace/foo, cell-43 for /workspace/bar). I'll keep an eye on their progress notifications."
        }
      ]
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script Running. Finish 1 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_2",
      "output": "Script Running. Finish 2 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script Running. Finish 3 tasks"
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Do you know what 1 + 1 equals?"
        }
      ]
    },
    {
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "1 + 1 = 2"
        }
      ]
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script Running. Finish 13 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_2",
      "output": "Script Running. Finish 23 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script Running. Finish 22 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script Running. Finish 24 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_2",
      "output": "Script Running. Finish 30 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script Running. Finish 31 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_2",
      "output": "Script Running. Finish 41 tasks"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_2",
      "output": "Script completed\nWall time 187.3 seconds\nOutput:\n44 passed"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script completed\nWall time 187.3 seconds\nOutput:\n42 passed"
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Have the backend tests finished running? What were the results? How many tests were successfully completed for each project? What is the task progress information, and who was notified for each of the two tasks?"
        }
      ]
    }
  ],
  "tool_choice": "auto",
  "parallel_tool_calls": false,
  "reasoning": {
    "context": "all_turns",
    "effort": "max",
    "summary": "detailed"
  },
  "store": false,
  "stream": true,
  "include": [
    "reasoning.encrypted_content"
  ]
}

Response:

{
    "type": "message",
    "status": "completed",
    "content": [
        {
            "type": "output_text",
            "annotations": [],
            "logprobs": [],
            "text": "Yes, both backend test suites finished successfully.\n\n- **Project `/workspace/foo`**: **44 passed** in 186.9 seconds.\n- **Project `/workspace/bar`**: **42 passed** in 186.9 seconds.\n- **Total**: **86 tests passed**, with no failures reported.\n\nTask progress messages indicated the tasks were still running at several checkpoints, including:\n\n- Finish 1 task\n- Finish 2 tasks\n- Finish 3 tasks\n- Finish 13 tasks\n- Finish 22 tasks\n- Finish 23 tasks\n- Finish 24 tasks\n- Finish 30 tasks\n- Finish 31 tasks\n- Finish 42 tasks\n\nBoth tasks ultimately completed. No information was provided indicating that anyone was notified for either task."
        }
    ],
    "phase": "final_answer",
    "role": "assistant"
}
Expected Demo

The same scenario with the proposal applied: the 12 late outputs (progress and completion) are delivered as user messages carrying the <notify cell_id="..." call_id="..."> envelope; the two yield outputs stay 1:1 tool outputs. Request:

{
  "model": "gpt-5.6-sol",
  "input": [
    {
      "type": "additional_tools",
      "role": "developer",
      "tools": [
        {
          "type": "custom",
          "name": "exec",
          "description": "Run JavaScript code to orchestrate/compose tool calls\n- Evaluates the provided JavaScript code in a fresh V8 isolate as an async module.\n- All nested tools are available on the global `tools` object, for example `await tools.exec_command(...)`.\n- Nested tool methods take either a string or an object as their input argument.\n- Nested tools return either an object or a string, based on the description.\n- Runs raw JavaScript -- no Node, no file system, no network access, no console.\n- You may optionally start the tool input with a first-line pragma like `// @exec: {\"yield_time_ms\": 10000, \"max_output_tokens\": 1000}`.\n- `yield_time_ms` asks `exec` to yield early if the script is still running. Defaults to 10000 ms. When it yields, the result says `Script running with cell ID ...` and the script keeps running in the background.\n- Global helpers:\n- `text(value: any)`: Appends a text item to the output.\n- `notify(value: any)`: immediately injects an extra `custom_tool_call_output` for the current `exec` call. Use it to report progress from a long-running script.\n- `store(key: string, value: any)` / `load(key: string)`: persist values between `exec` calls in the same session.\n- `setTimeout(callback: () => void, delayMs?: number)`: schedules a callback; pending timeouts do not keep `exec` alive by themselves.\n- `yield_control()`: yields the accumulated output to the model immediately while the script keeps running.\n\nNested tools:\n- `exec_command(args: {cmd: string, cwd?: string, yield_time_ms?: number, max_output_tokens?: number}): Promise<string>` -- Execute a shell command and return its output.",
          "format": {
            "type": "grammar",
            "syntax": "lark",
            "definition": "start: /(.|\\n)*/"
          }
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Run the test suites of both /workspace/foo and /workspace/bar for me at the same time, and keep me posted on the progress."
        }
      ]
    },
    {
      "type": "custom_tool_call",
      "call_id": "call_1",
      "name": "exec",
      "input": "// @exec: {\"yield_time_ms\": 1000}\nconst t0 = Date.now();\nconst list = await tools.exec_command({cmd: \"ls tests/test_*.py\", cwd: \"/workspace/foo\"});\nconst files = list.trim().split(\"\\n\");\nlet passed = 0;\nfor (const f of files) {\n  const out = await tools.exec_command({cmd: `pytest -q ${f}`, cwd: \"/workspace/foo\", yield_time_ms: 60000});\n  passed += Number(out.match(/(\\d+) passed/)?.[1] ?? 0);\n  notify(`Script Running. Finish ${passed} tasks`);\n}\nconst finalOut = await tools.exec_command({cmd: \"pytest -q tests/\", cwd: \"/workspace/foo\", yield_time_ms: 120000});\nconst total = Number(finalOut.match(/(\\d+) passed/)?.[1] ?? passed);\nconst secs = ((Date.now() - t0) / 1000).toFixed(1);\nnotify(`Script completed\\nWall time ${secs} seconds\\nOutput:\\n${total} passed`);"
    },
    {
      "type": "custom_tool_call",
      "call_id": "call_2",
      "name": "exec",
      "input": "// @exec: {\"yield_time_ms\": 1000}\nconst t0 = Date.now();\nconst list = await tools.exec_command({cmd: \"ls tests/test_*.py\", cwd: \"/workspace/bar\"});\nconst files = list.trim().split(\"\\n\");\nlet passed = 0;\nfor (const f of files) {\n  const out = await tools.exec_command({cmd: `pytest -q ${f}`, cwd: \"/workspace/bar\", yield_time_ms: 60000});\n  passed += Number(out.match(/(\\d+) passed/)?.[1] ?? 0);\n  notify(`Script Running. Finish ${passed} tasks`);\n}\nconst finalOut = await tools.exec_command({cmd: \"pytest -q tests/\", cwd: \"/workspace/bar\", yield_time_ms: 120000});\nconst total = Number(finalOut.match(/(\\d+) passed/)?.[1] ?? passed);\nconst secs = ((Date.now() - t0) / 1000).toFixed(1);\nnotify(`Script completed\\nWall time ${secs} seconds\\nOutput:\\n${total} passed`);"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_1",
      "output": "Script running with cell ID cell-42\nWall time 1.0 seconds\nOutput:\n(collecting...)"
    },
    {
      "type": "custom_tool_call_output",
      "call_id": "call_2",
      "output": "Script running with cell ID cell-43\nWall time 1.0 seconds\nOutput:\n(collecting...)"
    },
    {
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "Both test suites are now running in the background (cell-42 for /workspace/foo, cell-43 for /workspace/bar). I'll keep an eye on their progress notifications."
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-42\">Script Running. Finish 1 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-43\">Script Running. Finish 2 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-42\">Script Running. Finish 3 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Do you know what 1 + 1 equals?"
        }
      ]
    },
    {
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "1 + 1 = 2"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-42\">Script Running. Finish 13 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-43\">Script Running. Finish 23 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-42\">Script Running. Finish 22 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-42\">Script Running. Finish 24 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-43\">Script Running. Finish 30 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-42\">Script Running. Finish 31 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-43\">Script Running. Finish 41 tasks</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-43\">Script completed\nWall time 187.3 seconds\nOutput:\n44 passed</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "<notify cell_id=\"cell-42\">Script completed\nWall time 187.3 seconds\nOutput:\n42 passed</notify>"
        }
      ]
    },
    {
      "type": "message",
      "role": "user",
      "content": [
        {
          "type": "input_text",
          "text": "Have the backend tests finished running? What were the results? How many tests were successfully completed for each project? What is the task progress information, and who was notified for each of the two tasks?"
        }
      ]
    }
  ],
  "tool_choice": "auto",
  "parallel_tool_calls": false,
  "reasoning": {
    "context": "all_turns",
    "effort": "max",
    "summary": "detailed"
  },
  "store": false,
  "stream": true,
  "include": [
    "reasoning.encrypted_content"
  ]
}

Response (actual run against the same model, with the proposed shape):

{
    "type": "message",
    "status": "completed",
    "content": [
        {
            "type": "output_text",
            "annotations": [],
            "logprobs": [],
            "text": "Yes—both test runs finished successfully.\n\n| Project | Task/cell | Progress notifications observed | Final result |\n|---|---|---|---|\n| `/workspace/foo` | `cell-42` | `Finish 1`, `3`, `13`, `22`, `24`, and `31 tasks` | **42 passed** |\n| `/workspace/bar` | `cell-43` | `Finish 2`, `23`, `30`, and `41 tasks` | **44 passed** |\n\nBoth completion notifications reported **187.3 seconds** of wall time. No failures were reported in the final results, for **86 passed tests total** across the two projects.\n\nThe progress updates were interleaved because the suites ran concurrently. The notifications for both tasks were delivered to **you in this conversation**; no separate recipient or named person was specified."
        }
    ],
    "phase": "final_answer",
    "role": "assistant"
}

The model now attributes everything correctly — /workspace/foo: 42 passed, /workspace/bar: 44 passed — and even reports the progress notifications per task, which the buggy shape had merged into one undifferentiated list.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in codex-rs/core/src/tools/code_mode/delegate.rs around the notify() output construction, then read codex-rs/core/src/tools/code_mode/mod.rs around the cell-ID yield output. Reproduce the Bug Demo and Expected Demo to understand attribution across parallel cells. Done means late notify updates carry their cell association in message content, each exec call retains one tool output, and the 1:1 tool paths remain unchanged.

Written by the indexing model from the issue text.

Assessment

Tech stack
javascript, rust
Domain
backend-api-design, cli
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.