[QST] Why the pipeline wait before update tma descriptors are removed, at tma array gemm kernel
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 10.5k
- Forks
- 2.1k
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 7
Description
What is your question?
Before cutlass 4.0.0, there is a pipeline wait before update the tma descriptor (https://github.com/NVIDIA/cutlass/blob/2b78c2fe31d4adb4770f3ca226a7b4acc4e85e2b/include/cutlass/gemm/kernel/sm90_gemm_array_tma_warpspecialized_cooperative.hpp#L709-L713).
As commented:
// Purpose of this pipeline state is to make sure TMA loads have finished before doing descriptor updates
// Since this state is waiting for loads to finish, it must start in the inverted phase.
why it can be safely removed?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the referenced pre-4.0.0 implementation in include/cutlass/gemm/kernel/sm90_gemm_array_tma_warpspecialized_cooperative.hpp around lines 709-713, then compare the corresponding current TMA array GEMM kernel. Trace the pipeline wait and descriptor-update ordering, and document the evidence that establishes whether removing the wait is safe.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- backend, performance
- Issue type
- Refactor
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100