antirez / antirez/ds4

GLM 5.3 Flash: Metal resident-Q2 prefill −6–8% from 0e9cc2d (attention masking fix), measured on M5 Max

Abierto
#1,029 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
C
Estrellas
22.3k
Forks
2.1k
Merge medio
1 d 3 h
PR fusionados (30 d)
4

Descripción

**Setup:** MacBook Pro M5 Max 128 GB, `GLM-5.3-Flash-Q2.gguf` (resident, `--warm-weights`),
`ds4-server --ctx 393216 --vision GLM-5.3-Flash-Vision-Encoder.gguf --kv-disk-dir ... --kv-cache-cold-max-tokens 393216`,
server chat path, cold prefill on unique text (KV purged per block), server-reported `avg=`.

**Observation:** updating from `110afdd` to `6289c51` cost GLM cold prefill consistently at every depth,
while decode stayed flat. Interleaved A/B (3 rounds per condition, per-block checkout+binary+KV purge,
thermal settle + cooldowns; first round discarded as cold):

| depth | 110afdd | 6289c51 | delta |
| --- | ---: | ---: | ---: |
| 8K cold | 411 t/s | 381 t/s | −7.3% |
| 16K cold | 385 t/s | 362 t/s | −6.0% |
| 24K cold | 372 t/s | 341 t/s | −8.4% |
| 2K delta on warm 14K | 313 t/s | 278 t/s | −11% |
| decode 0–24K | 24.2–31.6 t/s | 24.2–30.4 t/s | flat/noise |

**Attribution:** direct two-commit differential, `0e9cc2d^` (=9d9e129) vs `0e9cc2d`
("Fix GLM attention masking and reductions across GPU backends"), 2 interleaved rounds each,
same harness:

| depth | 9d9e129 | 0e9cc2d | delta |
| --- | ---: | ---: | ---: |
| 8K cold | 415 t/s | 382 t/s | −8.0% |
| 16K cold | 386 t/s | 359 t/s | −7.0% |
| 24K cold | 376 t/s | 355 t/s | −5.5% |

That accounts for essentially the whole update delta. The commit message records the Strix
continued-prefill improvement (21→68 t/s) and unchanged official scores, but not a Metal cost,
so this may be an unnoticed side effect of the added reduction-storage synchronization /
padded-selection handling on Apple GPUs rather than an accepted trade-off.

Happy to run patches or diagnostics on this M5 Max config. (Same machine/config as the
measurements in #897/#898.)

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.