Hey bro, would you consider supporting Qwen3.8-27B-BF16?
- Lingua principale
- C
- Stelle
- 22.3k
- Fork
- 2.1k
- Merge medio
- 1g 3h
- PR unite (30g)
- 4
Descrizione
I recently tested Qwen3.8-27B-BF16 on a Mac M5 Max 128G using omlx and llama.cpp, but it's a dense architecture, and token output is extremely slow, only 10t/s. I could feel it repeatedly moving data into memory. However, the memory bandwidth seems to be only 600G/s. Do you know of any good ways to make it faster? I tried DFlash2, but the improvement was minimal. Also, I want to say that Alibaba wasn't exaggerating this time; this model performs significantly better than other smaller models in programming capabilities and long-term tasks. I personally think it's almost production-ready, because it handled programming tasks that DS4 couldn't solve on the first try. I highly recommend giving it a try. I hope you can consider this solution in your busy schedule.
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.