NVIDIA / NVIDIA/CUDALibrarySamples

SplitK for multiblock_gemm in cuBLASdx

Open
#192 1 comment 0 reactions 1 assignee View on GitHub

@llukas is already working on this.

Since Jun 27, 2025.

cuBLASdx
Dominant language
Cuda
Stars
2.5k
Forks
478
PR merge metrics
No merged PRs in 30d

Description

Hello!

I am currently learning CUTLASS and cuBLASdx and I have a question. multiblock_gemm.cu only allows K that fits in smem. I believe it can be extended to larger K following the splitK pattern here, but I am not quite sure how to implement this, I would appreciate suggestions!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.