NVIDIA / NVIDIA/cutlass

[QST] Why does sm103 GB300 FP4 block-scaled gemm have hardcoded SW128?

Open
#3,015 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

? - Needs Triage inactive-30d inactive-90d question
Dominant language
C++
Stars
10.5k
Forks
2.1k
Avg merge
3d 11h
Merged PRs (30d)
7

Description

Hi experts :-)

While I was using cutlass profiler & auto-tuned kernels for FP4 block-scaled gemm,
I noticed that they have high SMEM bank conflicts.
I used 8192x8192x9216 and 9728x16384x9216 shapes for MxNxK. I think the selected best kernels were using MMA 256x256x96.

I wanted to try using different SMEM swizzle factors, just in case, but it seems to be hardcoded with SW128 for sm103.

Is there a reason why it's fixed to SW128?

Thank you!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No source file or test is named. Start by locating the sm103 FP4 block-scaled GEMM configurations used by the profiler for the stated shapes and trace where SW128 is selected; reproduce the profiler runs and inspect the reported SMEM bank conflicts. Done means documenting why SW128 is fixed or identifying the specific configuration change needed to evaluate other swizzle factors.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.