[Feature]: Add FP8 KV Cache support to Aiter MLA
Open
@carlushuang is already working on this.
Since Aug 28, 2025.
High Priority
- Dominant language
- Python
- Stars
- 565
- Forks
- 585
- Avg merge
- 3d 4h
- Merged PRs (30d)
- 366
Description
Suggestion Description
- Adding FP8 KV Cache [Optional] to
mla_decode_fwdinterface. - Per tensor scaling, with precomputed scaling numbers [by quantizer's calibration],
first. - Per tensor per 128-channel (1x128), with dynamic-quant computed tile-based scaling numbers,
second - Support FP8 KV Cache transfer in PD for MLA decode
Operating System
No response
GPU
No response
ROCm Component
No response
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.