antirez / antirez/ds4

Hey bro, would you consider supporting Qwen3.8-27B-BF16?

Aperta
#866 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
C
Stelle
22.3k
Fork
2.1k
Merge medio
1g 3h
PR unite (30g)
4

Descrizione

I recently tested Qwen3.8-27B-BF16 on a Mac M5 Max 128G using omlx and llama.cpp, but it's a dense architecture, and token output is extremely slow, only 10t/s. I could feel it repeatedly moving data into memory. However, the memory bandwidth seems to be only 600G/s. Do you know of any good ways to make it faster? I tried DFlash2, but the improvement was minimal. Also, I want to say that Alibaba wasn't exaggerating this time; this model performs significantly better than other smaller models in programming capabilities and long-term tasks. I personally think it's almost production-ready, because it handled programming tasks that DS4 couldn't solve on the first try. I highly recommend giving it a try. I hope you can consider this solution in your busy schedule.

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.