Hey bro, would you consider supporting Qwen3.8-27B-BF16?
- Lenguaje dominante
- C
- Estrellas
- 22.3k
- Forks
- 2.1k
- Merge medio
- 1 d 3 h
- PR fusionados (30 d)
- 4
Descripción
I recently tested Qwen3.8-27B-BF16 on a Mac M5 Max 128G using omlx and llama.cpp, but it's a dense architecture, and token output is extremely slow, only 10t/s. I could feel it repeatedly moving data into memory. However, the memory bandwidth seems to be only 600G/s. Do you know of any good ways to make it faster? I tried DFlash2, but the improvement was minimal. Also, I want to say that Alibaba wasn't exaggerating this time; this model performs significantly better than other smaller models in programming capabilities and long-term tasks. I personally think it's almost production-ready, because it handled programming tasks that DS4 couldn't solve on the first try. I highly recommend giving it a try. I hope you can consider this solution in your busy schedule.
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.