antirez / antirez/ds4

Metal wires entire mmap-backed model views on first GPU use (ssd-streaming, mixed-size expert layers)

オープン
#638 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
C
スター
22.3k
フォーク
2.1k
平均マージ
1日 3時間
マージ済み PR(30日)
4

説明

On mixed-precision MoE models, layers whose experts fall outside the slab size class bypass the expert cache and read experts through mmap-backed model views wrapped in Metal buffers.

We measured (M4 Max 128 GB, macOS 26, DeepSeek 4 Flash: 6/43 such layers, 20.25 GiB of views) that the driver wires the **entire mapped views** at first decode use — 65.2 GiB wired vs 3.4 at rest — regardless of how few bytes each token actually touches. On 48 GB machines this makes any honest expert-cache budget impossible; at 64 GB it still consumes a third of the budget.

Isolated reproduction (~200-line C+Metal benchmark, same file, same 2 GiB of scattered reads under a GPU kernel):

- explicit 2 MiB-aligned slot buffers filled by `pread` and wrapped with `newBufferWithBytesNoCopy`: wired growth **≤ 0.8 GiB**;
- a single bytesNoCopy buffer over an 8 GiB mmap of the same file: wired growth **+8–10 GiB** for the same 2 GiB touched.

Verdict was identical across interleaved repeat pairs; checksums CPU=GPU=across-arms all runs. Happy to share the benchmark source.

Two independent public runtimes avoid this class of problem by never mapping the expert pool: TurboFieldfare (Swift/Metal, bounded slot pool + parallel `pread`) and WASTE (CPU, `pread` + 4 KiB records).

If there is interest, we can propose an opt-in slot-based read path for bypass-class layers, or contribute the measurement evidence to a design discussion. Related: the honest-measurement mode in #637 is what made the wired behavior visible in the first place.

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。