BYOK completions wire API fails with reasoning_content in streaming deltas (transient API error, 5 retries)
- Dominant language
- Shell
- Stars
- 11.2k
- Forks
- 1.9k
- Avg merge
- 14h 16m
- Merged PRs (30d)
- 6
Description
### Describe the bug
When using GitHub Copilot CLI with a BYOK provider that emits `reasoning_content` in streaming chat completion deltas (the `completions` wire API), Copilot reports **"Request failed due to a transient API error. Retrying..."** and retries 5 times before giving up — even though the provider returns HTTP 200 OK with valid SSE chunks.
The provider's streaming response includes the `reasoning_content` field in delta chunks (a de-facto standard extension used by reasoning models like DeepSeek, GLM, Qwen3, and surfaced via MLX servers like oMLX and Astronomical). Copilot's OpenAI completions client does not parse `reasoning_content` in the streaming delta, causing it to treat the response as a parse failure.
### Affected version
1.0.73 (Copilot CLI), macOS, darwin-arm64
### Steps to reproduce
1. Run a local OpenAI-compatible server that emits `reasoning_content` in streaming deltas for a reasoning model (e.g., Astronomical serving `mlx-community/Ornith-1.0-35B-OptiQ-4bit`, a Qwen-architecture reasoning model).
2. Configure BYOK:
```bash
export COPILOT_PROVIDER_BASE_URL=http://127.0.0.1:6732/v1
export COPILOT_PROVIDER_TYPE=openai
export COPILOT_PROVIDER_WIRE_API=completions
export COPILOT_MODEL=mlx-community/Ornith-1.0-35B-OptiQ-4bit
```
3. Launch `copilot` and send any message (e.g., `2+2?`).
4. Observe: "Request failed due to a transient API error. Retrying..." repeats 5 times.
### Expected behavior
Copilot CLI should parse `reasoning_content` in streaming delta chunks (as a reasoning delta) and surface it to the user — either via reasoning events or by ignoring it if reasoning display is not supported. It should not treat the response as a transient API error.
### What Copilot actually sends (captured via proxy)
```json
{
"model": "mlx-community/Ornith-1.0-35B-OptiQ-4bit",
"messages": [...],
"tools": [... 54 tools ...],
"stream": true,
"stream_options": {"include_usage": true}
}
```
### What the provider returns
```json
data: {"choices":[{"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"Here"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"'s"},"finish_reason":null}]}
...
data: {"choices":[{"delta":{"content":"4"},"finish_reason":null}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}],"usage":{...}}
data: [DONE]
```
The response is valid SSE with a proper `[DONE]` terminator and usage stats. HTTP status is 200. The only non-standard field is `reasoning_content` in the delta.
### Why this is a Copilot bug, not a provider bug
1. **`reasoning_content` is a de-facto standard** for reasoning models in the OpenAI-compatible ecosystem. It is recognized by models.dev (`"interleaved": {"field": "reasoning_content"}`) for models like GLM-5.1, GLM-5.2, and others.
2. **Other clients handle it fine.** OpenCode's OpenAI chat protocol explicitly parses `delta.reasoning_content` and emits it as a reasoning delta event (`packages/llm/src/protocols/openai-chat.ts` lines 419-420). OpenCode works perfectly with the same Astronomical provider.
3. **Other MLX servers emit it too.** oMLX (`omlx/api/thinking.py`, `omlx/api/adapters/openai.py`) and mlx-vlm both emit `reasoning_content` in streaming deltas for reasoning models. This is the standard way to surface `start_thinking...end_thinking` blocks in the OpenAI-compatible API.
4. **Related issue #3195** documents that Copilot CLI's reasoning handling for BYOK providers is incomplete — but that issue describes missing events, whereas this issue is more severe: the entire request fails and retries 5 times.
### Proposed fix
Copilot CLI's OpenAI completions streaming parser should:
1. Recognize `delta.reasoning_content` as a reasoning delta (not fail to parse it).
2. Emit it as a reasoning event (or silently ignore it if reasoning display is unsupported).
3. Continue processing `delta.content` normally.
At minimum, an unrecognized field in the delta object should not cause the entire response to be treated as a transient API error.
### Additional context
- Copilot CLI version: 1.0.73
- BYOK provider: Astronomical (local MLX server, OpenAI-compatible at `http://127.0.0.1:6732/v1`)
- Model: `mlx-community/Ornith-1.0-35B-OptiQ-4bit` (Qwen-architecture reasoning model)
- Wire API: `completions` (OpenAI Chat Completions at `/v1/chat/completions`)
- The Anthropic wire type (`COPILOT_PROVIDER_TYPE=anthropic`) does not have this issue because Copilot's Anthropic client natively handles `thinking` blocks.
- Related: #3195 (reasoning events not triggered for BYOK), #3196 (Responses API reasoning events empty)
Contributor guide
Research direction
Start with the BYOK completions streaming path and reproduce the failure using the provided Astronomical or another SSE provider that emits reasoning_content. Compare the expected handling with packages/llm/src/protocols/openai-chat.ts, which the issue cites as a working reference. Done means reasoning_content no longer triggers retries, delta.content is still processed, and the response completes normally.
Written by the indexing model from the issue text.
Assessment
- Domain
- api, backend
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 55/100