NVIDIA / NVIDIA/CUDALibrarySamples
Example 11 Performant Gemm cublasdx
Open
@llukas is already working on this.
Since Jul 4, 2025.
cuBLASdx
- Dominant language
- Cuda
- Stars
- 2.5k
- Forks
- 478
- PR merge metrics
- No merged PRs in 30d
Description
Hello I am trying to run the Cublasdx gemm on a gh200 which cta size of 128,128,32 on problem size (8192,8192,8192) and noticed that it is running at 1/3rd the speed of cublaslt with half, half input fp32 accumulator and was wondering if this was expected for the out of box example, or I am doing something wrong when running it? I am sorry if this is a novice question, really excited about the interface of this though.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.