deepseek-ai / deepseek-ai/DeepEP
How to decide the quant group scale
Open
- Dominant language
- Cuda
- Stars
- 10.1k
- Forks
- 1.4k
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 2
Description
Hi there, I am wondering why did you decide the quantization group scale is 128 BF16s? Is it just because the thread number fits 128 well, or you also thought about quantization precision? Thank you!
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.