antirez / antirez/ds4

MoE disk offload

Abierto
#202 3 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
C
Estrellas
22.4k
Forks
2.1k
Merge medio
2 d 13 h
PR fusionados (30 d)
5

Descripción

Hi! I've been working on MoE disk offload for llama.cpp and got some promising [results](https://github.com/ggml-org/llama.cpp/discussions/23324). Managed to run a model nearly 2x larger than available RAM at ~9 tok/s.

The approach is pretty simple: keep N expert slots in memory and page missing ones from GGUF via pread with LRU eviction. The natural integration point in ds4 would be `tensor_expert_bytes()`, instead of returning a raw mmap pointer, check an LRU slot pool first and pread on miss. Don't have a 96GB machine to test myself, but thought you might be interested.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.