antirez / antirez/ds4

GLM-5.3-Flash FP8 Optimized Inference Support

Đang mở
#1,015 3 bình luận 1 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.3k
Fork
2.1k
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
4

Mô tả

As stated in `docs/MODELS.md` (line 46):
```
| `glm53-fp8` | 305 GiB | Packaged native weights only; inference not implemented
```
inference for `GLM-5.3-Flash` FP8 model is not yet supported in the current `ds4` engine.

I tried simply patching `ds4.c` by routing GLM Q8_0 experts through the existing generic Metal ID kernels and adding Q8_0 coverage to the F32 reference path (see [https://github.com/zzyu17/ds4/commit/61b0d30c30d05cd118335f60e044a8f772225c01](url) for details), which gives the below results for a GLM-5.3-Flash Q8_0 GGUF on M3 Ultra 512G.

Image

The performance is quite good, with ~390 token/s prefill and ~21 token/s generation. However, an optimized (Metal) inference kernel for GLM-5.3-Flash FP8 model should perform even better, compared to the generic Metal kernel patch.

So, looking for the optimized (Metal) inference support for GLM-5.3-Flash FP8 model, eagerly!

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.