[QST] How to Interpret Canonical Layout Shapes
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 10.5k
- Forks
- 2.1k
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 7
Description
Hi,
I am trying to understand the GmmaDescriptor layouts in SM90. The code comments describe a complex interleaved layout, for example: LayoutType::B128 : Swizzle<3,4,3> o smem_ptr o ((8,m),(T,2)):((8T,SBO),(1, T )) . This also mentioned in the SM90 ptx documentation for TMA (https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#asynchronous-warpgroup-level-canonical-layouts)
However, when I use TMA(without cute) to load a tile with CU_TENSOR_MAP_SWIZZLE_128B and print the raw Shared Memory values, the data appears to be physically contiguous (Row/Col Major + Swizzle) and not interleaved by 8 rows.
So, how to Interpret Canonical Layout Shapes?
Thanks!
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the SM90 PTX documentation section on asynchronous warpgroup-level canonical layouts and the issue's LayoutType::B128 example. Compare that description with the reported CU_TENSOR_MAP_SWIZZLE_128B shared-memory output and clarify how the canonical shape relates to the physically observed layout. Done means the relationship between the interleaved representation and contiguous memory arrangement is clearly documented.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 28/100