NVIDIA / NVIDIA/CUDALibrarySamples

Example 11 Performant Gemm cublasdx

Open
#268 1 comment 0 reactions 1 assignee View on GitHub

@llukas is already working on this.

Since Jul 4, 2025.

cuBLASdx
Dominant language
Cuda
Stars
2.5k
Forks
478
PR merge metrics
No merged PRs in 30d

Description

Hello I am trying to run the Cublasdx gemm on a gh200 which cta size of 128,128,32 on problem size (8192,8192,8192) and noticed that it is running at 1/3rd the speed of cublaslt with half, half input fp32 accumulator and was wondering if this was expected for the out of box example, or I am doing something wrong when running it? I am sorry if this is a novice question, really excited about the interface of this though.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.