antirez / antirez/ds4

Hey bro, would you consider supporting Qwen3.8-27B-BF16?

Đang mở
#866 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.3k
Fork
2.1k
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
4

Mô tả

I recently tested Qwen3.8-27B-BF16 on a Mac M5 Max 128G using omlx and llama.cpp, but it's a dense architecture, and token output is extremely slow, only 10t/s. I could feel it repeatedly moving data into memory. However, the memory bandwidth seems to be only 600G/s. Do you know of any good ways to make it faster? I tried DFlash2, but the improvement was minimal. Also, I want to say that Alibaba wasn't exaggerating this time; this model performs significantly better than other smaller models in programming capabilities and long-term tasks. I personally think it's almost production-ready, because it handled programming tasks that DS4 couldn't solve on the first try. I highly recommend giving it a try. I hope you can consider this solution in your busy schedule.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.