BYOK completions wire API fails with reasoning_content in streaming deltas (transient API error, 5 retries)
- Lingua principale
- Shell
- Stelle
- 11.2k
- Fork
- 1.9k
- Merge medio
- 14h 16m
- PR unite (30g)
- 6
Descrizione
### Describe the bug
When using GitHub Copilot CLI with a BYOK provider that emits `reasoning_content` in streaming chat completion deltas (the `completions` wire API), Copilot reports **"Request failed due to a transient API error. Retrying..."** and retries 5 times before giving up — even though the provider returns HTTP 200 OK with valid SSE chunks.
The provider's streaming response includes the `reasoning_content` field in delta chunks (a de-facto standard extension used by reasoning models like DeepSeek, GLM, Qwen3, and surfaced via MLX servers like oMLX and Astronomical). Copilot's OpenAI completions client does not parse `reasoning_content` in the streaming delta, causing it to treat the response as a parse failure.
### Affected version
1.0.73 (Copilot CLI), macOS, darwin-arm64
### Steps to reproduce
1. Run a local OpenAI-compatible server that emits `reasoning_content` in streaming deltas for a reasoning model (e.g., Astronomical serving `mlx-community/Ornith-1.0-35B-OptiQ-4bit`, a Qwen-architecture reasoning model).
2. Configure BYOK:
```bash
export COPILOT_PROVIDER_BASE_URL=http://127.0.0.1:6732/v1
export COPILOT_PROVIDER_TYPE=openai
export COPILOT_PROVIDER_WIRE_API=completions
export COPILOT_MODEL=mlx-community/Ornith-1.0-35B-OptiQ-4bit
```
3. Launch `copilot` and send any message (e.g., `2+2?`).
4. Observe: "Request failed due to a transient API error. Retrying..." repeats 5 times.
### Expected behavior
Copilot CLI should parse `reasoning_content` in streaming delta chunks (as a reasoning delta) and surface it to the user — either via reasoning events or by ignoring it if reasoning display is not supported. It should not treat the response as a transient API error.
### What Copilot actually sends (captured via proxy)
```json
{
"model": "mlx-community/Ornith-1.0-35B-OptiQ-4bit",
"messages": [...],
"tools": [... 54 tools ...],
"stream": true,
"stream_options": {"include_usage": true}
}
```
### What the provider returns
```json
data: {"choices":[{"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"Here"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"'s"},"finish_reason":null}]}
...
data: {"choices":[{"delta":{"content":"4"},"finish_reason":null}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}],"usage":{...}}
data: [DONE]
```
The response is valid SSE with a proper `[DONE]` terminator and usage stats. HTTP status is 200. The only non-standard field is `reasoning_content` in the delta.
### Why this is a Copilot bug, not a provider bug
1. **`reasoning_content` is a de-facto standard** for reasoning models in the OpenAI-compatible ecosystem. It is recognized by models.dev (`"interleaved": {"field": "reasoning_content"}`) for models like GLM-5.1, GLM-5.2, and others.
2. **Other clients handle it fine.** OpenCode's OpenAI chat protocol explicitly parses `delta.reasoning_content` and emits it as a reasoning delta event (`packages/llm/src/protocols/openai-chat.ts` lines 419-420). OpenCode works perfectly with the same Astronomical provider.
3. **Other MLX servers emit it too.** oMLX (`omlx/api/thinking.py`, `omlx/api/adapters/openai.py`) and mlx-vlm both emit `reasoning_content` in streaming deltas for reasoning models. This is the standard way to surface `start_thinking...end_thinking` blocks in the OpenAI-compatible API.
4. **Related issue #3195** documents that Copilot CLI's reasoning handling for BYOK providers is incomplete — but that issue describes missing events, whereas this issue is more severe: the entire request fails and retries 5 times.
### Proposed fix
Copilot CLI's OpenAI completions streaming parser should:
1. Recognize `delta.reasoning_content` as a reasoning delta (not fail to parse it).
2. Emit it as a reasoning event (or silently ignore it if reasoning display is unsupported).
3. Continue processing `delta.content` normally.
At minimum, an unrecognized field in the delta object should not cause the entire response to be treated as a transient API error.
### Additional context
- Copilot CLI version: 1.0.73
- BYOK provider: Astronomical (local MLX server, OpenAI-compatible at `http://127.0.0.1:6732/v1`)
- Model: `mlx-community/Ornith-1.0-35B-OptiQ-4bit` (Qwen-architecture reasoning model)
- Wire API: `completions` (OpenAI Chat Completions at `/v1/chat/completions`)
- The Anthropic wire type (`COPILOT_PROVIDER_TYPE=anthropic`) does not have this issue because Copilot's Anthropic client natively handles `thinking` blocks.
- Related: #3195 (reasoning events not triggered for BYOK), #3196 (Responses API reasoning events empty)
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia con il percorso di streaming delle completions BYOK e riproduci il problema utilizzando il provider Astronomical fornito o un altro provider SSE che emetta reasoning_content. Confronta la gestione prevista con packages/llm/src/protocols/openai-chat.ts, che l’issue cita come riferimento funzionante. Il lavoro è completato quando reasoning_content non attiva più retry, delta.content continua a essere elaborato e la risposta termina normalmente.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Ambito
- api, backend
- Tipo di issue
- Bug
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Stato di attività
- Tranquilla
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 55/100