antirez / antirez/ds4

MoE disk offload

Offen
#202 3 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
C
Sterne
22.4k
Forks
2.1k
Ø Merge
2 T. 13 Std.
Gemergte PRs (30 T.)
5

Beschreibung

Hi! I've been working on MoE disk offload for llama.cpp and got some promising [results](https://github.com/ggml-org/llama.cpp/discussions/23324). Managed to run a model nearly 2x larger than available RAM at ~9 tok/s.

The approach is pretty simple: keep N expert slots in memory and page missing ones from GGUF via pread with LRU eviction. The natural integration point in ds4 would be `tensor_expert_bytes()`, instead of returning a raw mmap pointer, check an LRU slot pool first and pread on miss. Don't have a 96GB machine to test myself, but thought you might be interested.

Beitragsleitfaden

Beitragsleitfaden öffnen

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.