NVIDIA / NVIDIA/nvmath-python

Support 1 * 128 and 128 * 128 block-wise quant?

未关闭
#38 2 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

@szkarpinski 已经在做这个了。

开始于 2025年8月6日。

enhancement
主要语言
Cython
星标
601
派生
46
PR 合并指标
30 天内没有已合并 PR

描述

In the CUDA 12.9 cuBLASLt documentation, I noticed support for 1×128 and 128×128 block-wise quantization methods. However, I found that nvmath-python currently lacks bindings for this type of quantize approach. I wonder do we have any plan for support this approach?

https://docs.nvidia.com/cuda/cublas/index.html#cublasltmatmulmatrixscale-t

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。