MoE disk offload
Abierto
- Lenguaje dominante
- C
- Estrellas
- 22.4k
- Forks
- 2.1k
- Merge medio
- 2 d 13 h
- PR fusionados (30 d)
- 5
Descripción
Hi! I've been working on MoE disk offload for llama.cpp and got some promising [results](https://github.com/ggml-org/llama.cpp/discussions/23324). Managed to run a model nearly 2x larger than available RAM at ~9 tok/s.
The approach is pretty simple: keep N expert slots in memory and page missing ones from GGUF via pread with LRU eviction. The natural integration point in ds4 would be `tensor_expert_bytes()`, instead of returning a raw mmap pointer, check an LRU slot pool first and pread on miss. Don't have a 96GB machine to test myself, but thought you might be interested.
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.