Stale comment in `examples/cute/tutorial/hopper/wgmma_sm90.cu`
Open
Beginner friendly
Nobody has claimed this yet.
? - Needs Triage
CUTLASS C++
documentation
inactive-30d
- Dominant language
- C++
- Stars
- 10.5k
- Forks
- 2.1k
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 7
Description
The comment for the gemm_nt's tiled copyA/B says the thread layout used is 32x4, but the used thread layout is 16x8.
// Define the thread layouts (static)
TiledCopy copyA = make_tiled_copy(Copy_Atom<SM80_CP_ASYNC_CACHEALWAYS<uint128_t>, TA>{},
Layout<Shape<_16,_8>>{}, // Thr layout 32x4 m-major
Layout<Shape< _8,_1>>{});// Val layout 8x1 m-major
TiledCopy copyB = make_tiled_copy(Copy_Atom<SM80_CP_ASYNC_CACHEALWAYS<uint128_t>, TB>{},
Layout<Shape<_16,_8>>{}, // Thr layout 32x4 n-major
Layout<Shape< _8,_1>>{});// Val layout 8x1 n-major
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Open examples/cute/tutorial/hopper/wgmma_sm90.cu and inspect the gemm_nt tiled copyA and copyB definitions. Correct the stale thread-layout comments to match the shown 16x8 layouts, then review the surrounding example to ensure both comments are consistent. Done means the comments no longer say 32x4 while the code uses 16x8.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 90/100