Hey bro, would you consider supporting Qwen3.8-27B-BF16?
- Vorherrschende Sprache
- C
- Sterne
- 22.3k
- Forks
- 2.1k
- Ø Merge
- 1 T. 3 Std.
- Gemergte PRs (30 T.)
- 4
Beschreibung
I recently tested Qwen3.8-27B-BF16 on a Mac M5 Max 128G using omlx and llama.cpp, but it's a dense architecture, and token output is extremely slow, only 10t/s. I could feel it repeatedly moving data into memory. However, the memory bandwidth seems to be only 600G/s. Do you know of any good ways to make it faster? I tried DFlash2, but the improvement was minimal. Also, I want to say that Alibaba wasn't exaggerating this time; this model performs significantly better than other smaller models in programming capabilities and long-term tasks. I personally think it's almost production-ready, because it handled programming tasks that DS4 couldn't solve on the first try. I highly recommend giving it a try. I hope you can consider this solution in your busy schedule.
Beitragsleitfaden
Bewertung
Dieses Issue wurde noch nicht bewertet.