github / github/copilot-cli

BYOK completions wire API fails with reasoning_content in streaming deltas (transient API error, 5 retries)

Aberta
#4,196 2 comentários 0 reações 0 responsáveis Ver no GitHub
area:models
Linguagem predominante
Shell
Estrelas
11.2k
Forks
1.9k
Merge médio
14h 16min
PRs com merge (30d)
6

Descrição

### Describe the bug

When using GitHub Copilot CLI with a BYOK provider that emits `reasoning_content` in streaming chat completion deltas (the `completions` wire API), Copilot reports **"Request failed due to a transient API error. Retrying..."** and retries 5 times before giving up — even though the provider returns HTTP 200 OK with valid SSE chunks.

The provider's streaming response includes the `reasoning_content` field in delta chunks (a de-facto standard extension used by reasoning models like DeepSeek, GLM, Qwen3, and surfaced via MLX servers like oMLX and Astronomical). Copilot's OpenAI completions client does not parse `reasoning_content` in the streaming delta, causing it to treat the response as a parse failure.

### Affected version

1.0.73 (Copilot CLI), macOS, darwin-arm64

### Steps to reproduce

1. Run a local OpenAI-compatible server that emits `reasoning_content` in streaming deltas for a reasoning model (e.g., Astronomical serving `mlx-community/Ornith-1.0-35B-OptiQ-4bit`, a Qwen-architecture reasoning model).
2. Configure BYOK:
```bash
export COPILOT_PROVIDER_BASE_URL=http://127.0.0.1:6732/v1
export COPILOT_PROVIDER_TYPE=openai
export COPILOT_PROVIDER_WIRE_API=completions
export COPILOT_MODEL=mlx-community/Ornith-1.0-35B-OptiQ-4bit
```
3. Launch `copilot` and send any message (e.g., `2+2?`).
4. Observe: "Request failed due to a transient API error. Retrying..." repeats 5 times.

### Expected behavior

Copilot CLI should parse `reasoning_content` in streaming delta chunks (as a reasoning delta) and surface it to the user — either via reasoning events or by ignoring it if reasoning display is not supported. It should not treat the response as a transient API error.

### What Copilot actually sends (captured via proxy)

```json
{
"model": "mlx-community/Ornith-1.0-35B-OptiQ-4bit",
"messages": [...],
"tools": [... 54 tools ...],
"stream": true,
"stream_options": {"include_usage": true}
}
```

### What the provider returns

```json
data: {"choices":[{"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"Here"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"'s"},"finish_reason":null}]}
...
data: {"choices":[{"delta":{"content":"4"},"finish_reason":null}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}],"usage":{...}}
data: [DONE]
```

The response is valid SSE with a proper `[DONE]` terminator and usage stats. HTTP status is 200. The only non-standard field is `reasoning_content` in the delta.

### Why this is a Copilot bug, not a provider bug

1. **`reasoning_content` is a de-facto standard** for reasoning models in the OpenAI-compatible ecosystem. It is recognized by models.dev (`"interleaved": {"field": "reasoning_content"}`) for models like GLM-5.1, GLM-5.2, and others.

2. **Other clients handle it fine.** OpenCode's OpenAI chat protocol explicitly parses `delta.reasoning_content` and emits it as a reasoning delta event (`packages/llm/src/protocols/openai-chat.ts` lines 419-420). OpenCode works perfectly with the same Astronomical provider.

3. **Other MLX servers emit it too.** oMLX (`omlx/api/thinking.py`, `omlx/api/adapters/openai.py`) and mlx-vlm both emit `reasoning_content` in streaming deltas for reasoning models. This is the standard way to surface `start_thinking...end_thinking` blocks in the OpenAI-compatible API.

4. **Related issue #3195** documents that Copilot CLI's reasoning handling for BYOK providers is incomplete — but that issue describes missing events, whereas this issue is more severe: the entire request fails and retries 5 times.

### Proposed fix

Copilot CLI's OpenAI completions streaming parser should:
1. Recognize `delta.reasoning_content` as a reasoning delta (not fail to parse it).
2. Emit it as a reasoning event (or silently ignore it if reasoning display is unsupported).
3. Continue processing `delta.content` normally.

At minimum, an unrecognized field in the delta object should not cause the entire response to be treated as a transient API error.

### Additional context

- Copilot CLI version: 1.0.73
- BYOK provider: Astronomical (local MLX server, OpenAI-compatible at `http://127.0.0.1:6732/v1`)
- Model: `mlx-community/Ornith-1.0-35B-OptiQ-4bit` (Qwen-architecture reasoning model)
- Wire API: `completions` (OpenAI Chat Completions at `/v1/chat/completions`)
- The Anthropic wire type (`COPILOT_PROVIDER_TYPE=anthropic`) does not have this issue because Copilot's Anthropic client natively handles `thinking` blocks.
- Related: #3195 (reasoning events not triggered for BYOK), #3196 (Responses API reasoning events empty)

Guia de contribuição

Abrir o guia de contribuição

Direção de pesquisa

Comece pelo caminho de streaming de completions do BYOK e reproduza a falha usando o provedor Astronomical fornecido ou outro provedor SSE que emita reasoning_content. Compare o tratamento esperado com packages/llm/src/protocols/openai-chat.ts, que o issue cita como uma referência funcional. Está concluído quando reasoning_content não causar mais retries, delta.content continuar sendo processado e a resposta for concluída normalmente.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Domínio
api, backend
Tipo de issue
Bug
Dificuldade
3/5
Tempo estimado
1-2 dias
Status de atividade
Pouca atividade
Clareza
Razoavelmente clara
Facilidade para iniciantes
55/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.