antirez / antirez/ds4

GLM --mtp on Metal (server): greedy output diverges from sequential decode, identical requests become nondeterministic, ignore_eos is ignored

Abierto
#938 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
C
Estrellas
22.3k
Forks
2.1k
Merge medio
1 d 3 h
PR fusionados (30 d)
4

Descripción

## Setup

- Hardware: Mac Studio, M3 Ultra (80 GPU cores), 512 GB, Metal backend
- Models: `GLM-5.2-UD-Q2_K_RoutedQ2K.gguf` and full `GLM-5.3-UD-IQ2_XXS_RoutedIQ2XXS_blk78Q2K.gguf` (both rope models, `n_rot == 64`)
- Binary: upstream `main` at `110afdd` plus the one-line gate from PR #936. That fix is a prerequisite for measuring GLM MTP on these models at all — before it, the b0c31af dense-compact flash prefill (#932) corrupts every prompt with ≥ 24 tokens and masks everything below.
- Server: `ds4-server --mtp --mtp-timing`, fresh process with an empty `--kv-disk-dir` per configuration, single session, no SSD streaming
- Requests: OpenAI chat completions, `glm-*-chat` alias, `temperature: 0`, fixed seed, `max_tokens: 1024`, `ignore_eos: true`. Prompts are fixed slices of the tracked `rendered_prompts.txt` corpus: W1 ≈ 956 tokens, W2 ≈ 5836 tokens.

## Finding 1 — MTP greedy output diverges from sequential decode

Greedy speculative decoding must be token-exact against the sequential path. With identical prompts, seed, and sampling, all four comparisons produced different text:

| Model | Workload | no MTP | --mtp | greedy-identical |
|---|---|---:|---:|---|
| GLM-5.2 | W1 | 1024 tok | 415 tok | no |
| GLM-5.2 | W2 | 1024 tok | 105 tok | no |
| GLM-5.3 | W1 | 1024 tok | 81 tok | no |
| GLM-5.3 | W2 | 1024 tok | 1024 tok | no |

The GLM-5.3/W1 and GLM-5.2/W2 rows stop early (finding 3). The GLM-5.3/W2 row runs the full 1024 tokens on both sides and still diverges right after a position where the sequential path has `ignore_eos` suppress the EOS argmax — the MTP side neither commits EOS (it continues) nor picks the same next-best token, so the committed token at that position differs from both the raw argmax and the suppressed argmax.

## Finding 2 — identical requests become nondeterministic under MTP

On one GLM-5.2 + `--mtp` server, the same W1 request sent twice back to back:

- run 1: `gen=415 finish=stop`, text A
- run 2: `gen=222 finish=stop`, text B — same first ~300 chars, then diverges mid-prose (e.g. `"...al massimo delle FLOPS"` vs `"...al massimo delle capacità di calcolo"`)

The second request reuses the live session (rewind to the common prefix, final prompt token re-evaluated — server log shows a 1-token re-prefill). The same replay path without MTP is byte-identical across runs.

Two things we checked that do **not** explain it:

- `ds4_session_rewind()` does reset `glm_mtp_have`, `glm_mtp_rollback_valid`, and `mtp_draft_valid`, so a stale pending draft surviving the rewind is unlikely.
- #905 (unused suffix of a partially consumed MTP block left in the live KV) is adjacent, but its fix is not on `main` at the time of writing and its documented symptom is prefix invalidation, not output divergence. It may still be the underlying mechanism; we could not confirm.

## Finding 3 — `ignore_eos` is ignored on the GLM spec path

Three of four MTP runs ended with `finish=stop` on a stop token despite `ignore_eos: true`; all no-MTP runs reached the 1024-token cap. Code-level cause: the GLM branch of `ds4_session_eval_speculative_argmax_impl` discards the EOS state (`(void)eos_token;`) and `ds4_session_glm_spec_cycle` does not take EOS information at all. The first token still comes from the suppressing argmax, but the verified/drafted second token is a raw argmax — when that is a stop token it gets committed and the server loop stops, where the sequential `argmax_ignoring_eos` path would have suppressed it and continued.

Normal requests want to stop at EOS, so this mostly affects benchmark-style `ignore_eos` clients — but it also means MTP changes stopping behavior relative to the non-MTP server.

## Acceptance and throughput context (for #920 / #676)

`--mtp-timing` cycle lines over the same runs:

- GLM-5.2: 296/446 cycles accepted (66.4%)
- GLM-5.3: 445/737 accepted (60.4%)

Decode throughput (server decode logs, tok/s):

| Model | region | no MTP | --mtp | delta |
|---|---|---:|---:|---:|
| GLM-5.3 | W1 dense | 16.76 | 13.32 / 13.52 | ≈ −20% |
| GLM-5.3 | W2 indexed | 15.97 | 8.83 | ≈ −45% |
| GLM-5.2 | W1 dense | 16.75 | 10.41 / 10.94 | ≈ −37% |
| GLM-5.2 | W2 indexed | 16.00 | 9.03 | ≈ −44% |

At 60–66% acceptance the width-2 cycle costs ≈ 1.8–2.2× a sequential step on this GPU, so MTP is a net loss for these rope models on M3 Ultra even before the correctness findings — consistent with the economics reported in #676 and the verify-cost work in #920 (which shows the opposite sign on M4 Max with GLM-5.3 Flash Q2, so this is model/hardware dependent).

## Repro sketch

```sh
./ds4-server --host 127.0.0.1 --port 8001 --mtp --mtp-timing \
--model gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf --ctx 393216 \
--kv-disk-dir /tmp/fresh-kv --kv-disk-space-mb 2048

# same request twice, compare responses (finding 2), then once more
# without --mtp for the baseline (findings 1 and 3)
curl -s localhost:8001/v1/chat/completions -d '{
"model": "glm-5.2-chat",
"messages": [{"role":"user","content": "<~950 tokens of fixed corpus text>\n\nContinue this document."}],
"temperature": 0, "seed": 1, "max_tokens": 1024, "ignore_eos": true
}'
```

## Related

- #932 / PR #936 — the b0c31af attention regression that previously masked all of the above on rope models
- #905 — unused MTP block suffix in the live KV; same family as finding 2, fix not on `main` yet
- #920 — width-2 verify cost on Metal (throughput context)
- #676 — speculative decoding net-negative economics on Apple silicon (DSpark)

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.