antirez / antirez/ds4

Proposal: serve token_embd.weight row lookups from the host mapping (~1 GiB residency win on streaming setups)

Abierto
#688 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
C
Estrellas
22.3k
Forks
2.1k
Merge medio
1 d 3 h
PR fusionados (30 d)
4

Descripción

### Observation

`token_embd.weight` is accessed one row (~8-16 KB) per decoded token, but it is prepared as a fully resident device span at startup -- ~1 GiB on DeepSeek V4 Flash Q8_0. On streaming setups (the configurations `--ssd-streaming` exists for), it is the single largest always-resident tensor, and its access pattern -- one read-only row per token -- is exactly what a read-only mmap serves at page-cache speed.

### Proposal

Under an env gate (`DS4_EMBD_MMAP=1`, default OFF): leave `token_embd.weight` out of the resident device spans at startup, and have the embed path resolve it to the read-only host model mapping instead. Hot rows stay in page cache; per-token decode touches one row.

### Measurements (GB10, DeepSeek V4 Flash Q8_0, streaming, ctx 131072)

| measure | resident (baseline) | DS4_EMBD_MMAP=1 |
|---|---|---|
| residency | 0.99 GiB device spans | 0.00 GiB -- **-1010 MiB resident** |
| warm decode e2e t/s (3 turns) | 4.25 / 4.28 / 4.36 | 4.19 / 4.23 / 4.33 (-0.9% avg, within run-to-run noise) |
| cold prefill (3430 tok) | baseline | no cost signal |
| worst case: first token after a full page-cache drop | -- | one-time ~2.5 s transient (4.8 s vs 2.3 s warm), recovered by the next request |

Verdict on that box: ~1 GiB of residency freed for no measurable throughput cost.

### Important caveat

The measurement and the patch rely on **unified-memory ATS/HMM pageable access** (GB10 / Grace class): the GPU reads the file-backed host pages directly. On discrete-GPU targets the same idea would need the `cudaHostRegister`'d-mapping path that `cuda_model_range_ptr` already has -- untested by me, and honestly future work. The env gate defaults OFF precisely so non-unified targets are unaffected.

PR with the env-gated implementation attached.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.