[QST]What is the best layout for GEMM?
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 10.5k
- Forks
- 2.1k
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 7
Description
I'm learning the example codes for tensor op gemm.
In some examples, the A, B, C are all column major.
Why the A B C are all column major? Is this the best layout on GPU for gemm?
I read the flash attention implementation, when it computes Q@K, it put Q as row major, and K is column major.
But why here we use A B both in column major? How to select the A B layouts on different gpus?
Any one give some explanation? Thank you!
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No file, test, or entry point is identified; start by locating the tensor-op GEMM examples that use column-major A, B, and C, then compare them with the flash-attention Q@K layout mentioned here. Done means explaining the layout choices and how to select layouts across different GPUs, with the explanation added to an appropriate documentation location.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- hpc, performance
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100