[QST] How to implement a fused mixed precision matrix multiplication such as w4a4 + w16a16?
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 10.5k
- Forks
- 2.1k
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 7
Description
Dear Team,
I wish to implement a fused mixed precision matrix multiplication such as w4a4 + w16a16 where the w16a16 part is small. An example of this kernel used is for accelerating an LLM with LoRA applied.
I can find some examples in "torchao" that implement matrix multiplication of w4a4/w4a8 and integrate matrix multiplication and dequantization via epilogue, but I don't know how to further integrate matrix multiplication of w16a16 on top of it, is there any examples I can refer to?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the torchao examples for w4a4 and w4a8, especially how matrix multiplication and dequantization are combined through an epilogue. Then investigate how a small w16a16 operation for the LoRA use case could be integrated. The issue names no CUTLASS files, tests, entry points, or concrete completion criterion.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp, python
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100