BYOK completions wire API fails with reasoning_content in streaming deltas (transient API error, 5 retries)
- 主要言語
- Shell
- スター
- 11.2k
- フォーク
- 1.9k
- 平均マージ
- 14時間 16分
- マージ済み PR(30日)
- 6
説明
### Describe the bug
When using GitHub Copilot CLI with a BYOK provider that emits `reasoning_content` in streaming chat completion deltas (the `completions` wire API), Copilot reports **"Request failed due to a transient API error. Retrying..."** and retries 5 times before giving up — even though the provider returns HTTP 200 OK with valid SSE chunks.
The provider's streaming response includes the `reasoning_content` field in delta chunks (a de-facto standard extension used by reasoning models like DeepSeek, GLM, Qwen3, and surfaced via MLX servers like oMLX and Astronomical). Copilot's OpenAI completions client does not parse `reasoning_content` in the streaming delta, causing it to treat the response as a parse failure.
### Affected version
1.0.73 (Copilot CLI), macOS, darwin-arm64
### Steps to reproduce
1. Run a local OpenAI-compatible server that emits `reasoning_content` in streaming deltas for a reasoning model (e.g., Astronomical serving `mlx-community/Ornith-1.0-35B-OptiQ-4bit`, a Qwen-architecture reasoning model).
2. Configure BYOK:
```bash
export COPILOT_PROVIDER_BASE_URL=http://127.0.0.1:6732/v1
export COPILOT_PROVIDER_TYPE=openai
export COPILOT_PROVIDER_WIRE_API=completions
export COPILOT_MODEL=mlx-community/Ornith-1.0-35B-OptiQ-4bit
```
3. Launch `copilot` and send any message (e.g., `2+2?`).
4. Observe: "Request failed due to a transient API error. Retrying..." repeats 5 times.
### Expected behavior
Copilot CLI should parse `reasoning_content` in streaming delta chunks (as a reasoning delta) and surface it to the user — either via reasoning events or by ignoring it if reasoning display is not supported. It should not treat the response as a transient API error.
### What Copilot actually sends (captured via proxy)
```json
{
"model": "mlx-community/Ornith-1.0-35B-OptiQ-4bit",
"messages": [...],
"tools": [... 54 tools ...],
"stream": true,
"stream_options": {"include_usage": true}
}
```
### What the provider returns
```json
data: {"choices":[{"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"Here"},"finish_reason":null}]}
data: {"choices":[{"delta":{"reasoning_content":"'s"},"finish_reason":null}]}
...
data: {"choices":[{"delta":{"content":"4"},"finish_reason":null}]}
data: {"choices":[{"delta":{},"finish_reason":"stop"}],"usage":{...}}
data: [DONE]
```
The response is valid SSE with a proper `[DONE]` terminator and usage stats. HTTP status is 200. The only non-standard field is `reasoning_content` in the delta.
### Why this is a Copilot bug, not a provider bug
1. **`reasoning_content` is a de-facto standard** for reasoning models in the OpenAI-compatible ecosystem. It is recognized by models.dev (`"interleaved": {"field": "reasoning_content"}`) for models like GLM-5.1, GLM-5.2, and others.
2. **Other clients handle it fine.** OpenCode's OpenAI chat protocol explicitly parses `delta.reasoning_content` and emits it as a reasoning delta event (`packages/llm/src/protocols/openai-chat.ts` lines 419-420). OpenCode works perfectly with the same Astronomical provider.
3. **Other MLX servers emit it too.** oMLX (`omlx/api/thinking.py`, `omlx/api/adapters/openai.py`) and mlx-vlm both emit `reasoning_content` in streaming deltas for reasoning models. This is the standard way to surface `start_thinking...end_thinking` blocks in the OpenAI-compatible API.
4. **Related issue #3195** documents that Copilot CLI's reasoning handling for BYOK providers is incomplete — but that issue describes missing events, whereas this issue is more severe: the entire request fails and retries 5 times.
### Proposed fix
Copilot CLI's OpenAI completions streaming parser should:
1. Recognize `delta.reasoning_content` as a reasoning delta (not fail to parse it).
2. Emit it as a reasoning event (or silently ignore it if reasoning display is unsupported).
3. Continue processing `delta.content` normally.
At minimum, an unrecognized field in the delta object should not cause the entire response to be treated as a transient API error.
### Additional context
- Copilot CLI version: 1.0.73
- BYOK provider: Astronomical (local MLX server, OpenAI-compatible at `http://127.0.0.1:6732/v1`)
- Model: `mlx-community/Ornith-1.0-35B-OptiQ-4bit` (Qwen-architecture reasoning model)
- Wire API: `completions` (OpenAI Chat Completions at `/v1/chat/completions`)
- The Anthropic wire type (`COPILOT_PROVIDER_TYPE=anthropic`) does not have this issue because Copilot's Anthropic client natively handles `thinking` blocks.
- Related: #3195 (reasoning events not triggered for BYOK), #3196 (Responses API reasoning events empty)
コントリビューションガイド
調査の方向性
BYOK の completions ストリーミングパスから始め、提供されている Astronomical または reasoning_content を送信する別の SSE プロバイダーを使って失敗を再現します。期待される処理を、issue が動作するリファレンスとして挙げている packages/llm/src/protocols/openai-chat.ts と比較します。reasoning_content によってリトライが発生しなくなり、delta.content が引き続き処理され、レスポンスが正常に完了すれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 領域
- api, backend
- issue の種類
- バグ
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 活発さ
- 静か
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 55/100