NVIDIA / NVIDIA/cutlass

[QST] Why the pipeline wait before update tma descriptors are removed, at tma array gemm kernel

Open
#2,912 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

? - Needs Triage inactive-30d inactive-90d question
Dominant language
C++
Stars
10.5k
Forks
2.1k
Avg merge
3d 11h
Merged PRs (30d)
7

Description

What is your question?

Before cutlass 4.0.0, there is a pipeline wait before update the tma descriptor (https://github.com/NVIDIA/cutlass/blob/2b78c2fe31d4adb4770f3ca226a7b4acc4e85e2b/include/cutlass/gemm/kernel/sm90_gemm_array_tma_warpspecialized_cooperative.hpp#L709-L713).

As commented:

// Purpose of this pipeline state is to make sure TMA loads have finished before doing descriptor updates
// Since this state is waiting for loads to finish, it must start in the inverted phase.

why it can be safely removed?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the referenced pre-4.0.0 implementation in include/cutlass/gemm/kernel/sm90_gemm_array_tma_warpspecialized_cooperative.hpp around lines 709-713, then compare the corresponding current TMA array GEMM kernel. Trace the pipeline wait and descriptor-update ordering, and document the evidence that establishes whether removing the wait is safe.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
backend, performance
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.