Support 1 * 128 and 128 * 128 block-wise quant?
未关闭
@szkarpinski 已经在做这个了。
开始于 2025年8月6日。
enhancement
- 主要语言
- Cython
- 星标
- 601
- 派生
- 46
- PR 合并指标
- 30 天内没有已合并 PR
描述
In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quantize approach. I wonder do we have any plan for support this approach?
https://docs.nvidia.com/cuda/cublas/index.html#cublasltmatmulmatrixscale-t
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
评估
这个 Issue 还没有评估数据。