BYOK completions wire API fails with reasoning_content in streaming deltas (transient API error, 5 retries)
- Linguagem predominante
- Shell
- Estrelas
- 11.2k
- Forks
- 1.9k
- Merge médio
- 14h 16min
- PRs com merge (30d)
- 6
Descrição
### Describe the bug
When using GitHub Copilot CLI with a BYOK provider that emits `reasoning_content` in streaming chat completion deltas (the `completions` wire API), Copilot reports **"Request failed due to a transient API error. Retrying..."** and retries 5 times before giving up — even though the provider returns HTTP 200 OK with valid SSE chunks.
The provider's streaming response includes the `reasoning_content` field in delta chunks (a de-facto standard extension used by reasoning models like DeepSeek, GLM, Qwen3, and surfaced via MLX servers like oMLX and Astronomical). Copilot's OpenAI completions client does not parse `reasoning_content` in the streaming delta, causing it to treat the response as a parse failure.
### Affected version
1.0.73 (Copilot CLI), macOS, darwin-arm64
### Steps to reproduce
1. Run a local OpenAI-compatible server that emits `reasoning_content` in streaming deltas for a reasoning model (e.g., Astronomical serving `mlx-community/Ornith-1.0-35B-OptiQ-4bit`, a Qwen-architecture reasoning model).
2. Configure BYOK:
```bash
export COPILOT_PROVIDER_BASE_URL=http://127.0.0.1:6732/v1
export COPILOT_PROVIDER_TYPE=openai
export COPILOT_PROVIDER_WIRE_API=completions
export COPILOT_MODEL=mlx-community/Ornith-1.0-35B-OptiQ-4bit
```
3. Launch `copilot` and send any message (e.g., `2+2?`).
4. Observe: "Request failed due to a transient API error. Retrying..." repeats 5 times.
### Expected behavior
Copilot CLI should parse `reasoning_content` in streaming delta chunks (as a reasoning delta) and surface it to the user — either via reasoning events or by ignoring it if reasoning display is not supported. It should not treat the response as a transient API error.
### What Copilot actually sends (captured via proxy)
```json
{
"model": "mlx-community/Ornith-1.0-35B-OptiQ-4bit",
"messages": [...],
"tools": [... 54 tools ...],
"stream": true,
"stream_options": {"include_usage": true}
}
```
### What the provider returns
```json
data: {"choices":[{"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"Here"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"'s"},"finish_reason":null}]}
...
data: {"choices":[{"delta":{"content":"4"},"finish_reason":null}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}],"usage":{...}}
data: [DONE]
```
The response is valid SSE with a proper `[DONE]` terminator and usage stats. HTTP status is 200. The only non-standard field is `reasoning_content` in the delta.
### Why this is a Copilot bug, not a provider bug
1. **`reasoning_content` is a de-facto standard** for reasoning models in the OpenAI-compatible ecosystem. It is recognized by models.dev (`"interleaved": {"field": "reasoning_content"}`) for models like GLM-5.1, GLM-5.2, and others.
2. **Other clients handle it fine.** OpenCode's OpenAI chat protocol explicitly parses `delta.reasoning_content` and emits it as a reasoning delta event (`packages/llm/src/protocols/openai-chat.ts` lines 419-420). OpenCode works perfectly with the same Astronomical provider.
3. **Other MLX servers emit it too.** oMLX (`omlx/api/thinking.py`, `omlx/api/adapters/openai.py`) and mlx-vlm both emit `reasoning_content` in streaming deltas for reasoning models. This is the standard way to surface `start_thinking...end_thinking` blocks in the OpenAI-compatible API.
4. **Related issue #3195** documents that Copilot CLI's reasoning handling for BYOK providers is incomplete — but that issue describes missing events, whereas this issue is more severe: the entire request fails and retries 5 times.
### Proposed fix
Copilot CLI's OpenAI completions streaming parser should:
1. Recognize `delta.reasoning_content` as a reasoning delta (not fail to parse it).
2. Emit it as a reasoning event (or silently ignore it if reasoning display is unsupported).
3. Continue processing `delta.content` normally.
At minimum, an unrecognized field in the delta object should not cause the entire response to be treated as a transient API error.
### Additional context
- Copilot CLI version: 1.0.73
- BYOK provider: Astronomical (local MLX server, OpenAI-compatible at `http://127.0.0.1:6732/v1`)
- Model: `mlx-community/Ornith-1.0-35B-OptiQ-4bit` (Qwen-architecture reasoning model)
- Wire API: `completions` (OpenAI Chat Completions at `/v1/chat/completions`)
- The Anthropic wire type (`COPILOT_PROVIDER_TYPE=anthropic`) does not have this issue because Copilot's Anthropic client natively handles `thinking` blocks.
- Related: #3195 (reasoning events not triggered for BYOK), #3196 (Responses API reasoning events empty)
Guia de contribuição
Direção de pesquisa
Comece pelo caminho de streaming de completions do BYOK e reproduza a falha usando o provedor Astronomical fornecido ou outro provedor SSE que emita reasoning_content. Compare o tratamento esperado com packages/llm/src/protocols/openai-chat.ts, que o issue cita como uma referência funcional. Está concluído quando reasoning_content não causar mais retries, delta.content continuar sendo processado e a resposta for concluída normalmente.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Domínio
- api, backend
- Tipo de issue
- Bug
- Dificuldade
- 3/5
- Tempo estimado
- 1-2 dias
- Status de atividade
- Pouca atividade
- Clareza
- Razoavelmente clara
- Facilidade para iniciantes
- 55/100