antirez / antirez/ds4

OpenAI /v1/chat/completions tool rounds never hit live continuation (match=tool-output-ids); live slot discarded despite full-prompt common

Đang mở
#609 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.3k
Fork
2.1k
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
4

Mô tả

On the Anthropic endpoint, tool rounds continue the live KV slot cheaply (`anthropic live continuation match=tool-output-ids`, delta-only prefill — see #500 for a working example at 85K context with 400–900 token deltas). On the OpenAI chat/completions endpoint, the equivalent never happens: every tool round logs a live miss where `common` equals the entire previous prompt, the live slot is discarded, and prefill restarts from the oldest disk checkpoint.

Example (Open WebUI, native function calling, OpenAI protocol — four consecutive tool rounds in one conversation):

```
live kv cache miss live=6034 prompt=6128 common=5817 reason=token-mismatch
kv cache stored tokens=6034 trimmed=0 reason=evict key=token-text size=102.08 MiB save=16.6 ms
chat ctx=5779..6128:349 TOOLS prompt done 1.857s
live kv cache miss live=6213 prompt=8238 common=6128 reason=token-mismatch
kv cache stored tokens=6213 trimmed=0 reason=evict key=token-text size=104.43 MiB save=16.5 ms
chat ctx=5779..8238:2459 TOOLS prompt done 9.231s
live kv cache miss live=8633 prompt=14903 common=8238 reason=token-mismatch
kv cache stored tokens=8633 trimmed=0 reason=evict key=token-text size=136.21 MiB save=15.1 ms
chat ctx=5779..14903:9124 TOOLS prompt done 31.253s
live kv cache miss live=15139 prompt=20460 common=14903 reason=token-mismatch
kv cache stored tokens=15139 trimmed=0 reason=evict key=token-text size=221.61 MiB save=25.7 ms
```

Note the pattern: `common` always reaches exactly the previous prompt's end (6128, 8238, 14903), because the only divergence is the re-serialised assistant tool-call turn — the client strips thinking to an empty `` and re-encodes the tool call in OpenAI JSON format. Instead of rolling back to `common` and prefilling the delta, the live slot is discarded and prefill restarts from the system-prompt disk checkpoint (5779), so each round pays (context − 5779) tokens. At ~90K context this compounds to ~6-minute tool rounds.

The evicted snapshots store the divergent generated tail untrimmed (`trimmed=0` — this persists with `--kv-cache-boundary-trim-tokens 1024`, which appears to only affect cold-path saves), so they accumulate with `hits=0` and evict each other; only the system-prompt checkpoint ever hits.

Exact DSML tool replay is active and reports healthy: tool-call lines show `raw_tool_text=1`, and `--trace` shows `tool_replay: mem=N disk=0 canonical=0 missing_ids=0` with correct `call_*` IDs present in the replayed request.

A rollback-to-common on the live slot, or an OpenAI-path equivalent of the `tool-output-ids` continuation matching on `tool_call_id`, would reduce these rounds to delta-only prefill like the Anthropic path.

**Environment**

- ds4: main @ 0a7ad77
- macOS 26.5.2, Apple M5 Max, 128 GB
- Model: `DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf` (fully resident, no SSD streaming)
- Launch: `./ds4-server -m --ctx 100000 --kv-disk-dir ~/.cache/ds4-kv --kv-disk-space-mb 8192 --kv-cache-boundary-trim-tokens 1024 --host 127.0.0.1 --port 8300`
- Client: Open WebUI 0.10.2, Native function calling, streaming, `stream_options.include_usage`

Happy to test a branch — this setup reproduces the pattern on every multi-round tool conversation.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.