CommandCodeAI / CommandCodeAI/command-code

Provider API: SSE stream ends silently mid-tool-call during large tool-call argument generation (no [DONE], no finish_reason)

Đang mở
#785 5 bình luận 0 reaction 1 người được giao Xem trên GitHub

@ahmadbilaldev đang làm issue này rồi.

Từ ngày 17/9/2026.

Ngôn ngữ chính
Không có dữ liệu ngôn ngữ
Star
4k
Fork
350
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

Summary

Streaming chat.completions against the Provider API intermittently ends the SSE stream mid-tool-call while the model is generating a large tool-call JSON object. The stream stops cleanly — no [DONE], no finish_reason chunk, no error — with the tool-call arguments still incomplete. A client cannot distinguish this from a network drop.

Correlated with:

  • Large tool-call argument payloads (~10k–50k chars, e.g. a write_file/terminal call carrying a long body)
  • Large prompt context (~50k+ tokens)

Small tool-call arguments, plain text output of any length, and the same request in non-streaming mode all complete reliably. This looks like a streaming-path-specific drop in the upstream SSE implementation, not a hard output-token cap.

Expected Behavior

The SSE stream should run to completion ([DONE], finish_reason: "tool_calls") with complete JSON tool-call arguments ending in }, regardless of argument size.

Actual Behavior

The stream ends after a few thousand chunks with no terminator of any kind — no [DONE], no finish_reason chunk. The tool-call arguments are truncated mid-string. Total delivered arguments ~20–50KB (roughly 5–9k completion tokens), far below the requested max_tokens.

Steps to reproduce the issue

Send a streaming chat completion that forces the model to emit one big tool call:

{
  "model": "deepseek/deepseek-v4-flash",
  "messages": [
    {
      "role": "user",
      "content": "Call write_file with path '/tmp/big.txt' and content being a 20,000-character essay about Vietnamese coffee culture. The ENTIRE content must be inside the tool call arguments. Do not abbreviate."
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "write_file",
        "description": "Write a file to disk",
        "parameters": {
          "type": "object",
          "properties": {
            "path": {"type": "string"},
            "content": {"type": "string"}
          },
          "required": ["path", "content"]
        }
      }
    }
  ],
  "tool_choice": {"type": "function", "function": {"name": "write_file"}},
  "max_tokens": 131072,
  "stream": true
}

Python repro (stdlib only):

import json, urllib.request, os

key = os.environ["COMMANDCODE_API_KEY"]
url = "https://api.commandcode.ai/provider/v1/chat/completions"
body = {
    "model": "deepseek/deepseek-v4-flash",
    "messages": [{"role": "user", "content": "Call write_file with path '/tmp/big.txt' and content being a 20,000-character essay about Vietnamese coffee culture. The ENTIRE content must be inside the tool call arguments. Do not abbreviate."}],
    "tools": [{"type": "function", "function": {"name": "write_file", "description": "Write a file to disk", "parameters": {"type": "object", "properties": {"path": {"type": "string"}, "content": {"type": "string"}}, "required": ["path", "content"]}}}],
    "tool_choice": {"type": "function", "function": {"name": "write_file"}},
    "max_tokens": 131072,
    "stream": True,
}
req = urllib.request.Request(url, data=json.dumps(body).encode(),
    headers={"Content-Type": "application/json", "Authorization": "Bearer " + key})
with urllib.request.urlopen(req, timeout=300) as r:
    chunks = 0
    for raw in r:
        line = raw.decode("utf-8", "replace").strip()
        if line.startswith("data: "):
            line = line[6:]
        if line == "[DONE]":
            break
        chunks += 1
        print(line[:200])
    print("chunks:", chunks)

Expected: stream continues to [DONE] with finish_reason: "tool_calls", complete JSON args ending in }.
Observed (when it fails): stream simply ends after a few thousand chunks — no [DONE], no finish_reason chunk, tool arguments truncated mid-string.

Command Code Version

Provider API (gateway) — api.commandcode.ai/provider/v1/chat/completions, model deepseek/deepseek-v4-flash

Operating System

Linux

Terminal/IDE

OpenAI-compatible client (raw SSE reader)

Shell

No response

Session file (optional)

No response

Fix prompt (optional)

No response

Additional context
  • The non-streaming equivalent of the same request completes fine — a 190KB tool-call payload via stream: false succeeds. So any cap is in the streaming path only.
  • The failure reproduces through a plain OpenAI-compatible client; no client SDK or router involved.
  • Frequency appears to increase with prompt context size; the smallest repro used an ~8k-token prompt.
  • This is distinct from #628 and #755 (both CLI-side handling of a mid-flight drop that already ended with an error). Here the stream ends silently with no terminator at all.

Happy to run additional repros (specific context sizes, different models in the catalog, non-SSE streaming) on request.

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.