ROCm / ROCm/aiter

[Feature]: Add FP8 KV Cache support to Aiter MLA

Open
#899 2 comments 1 reaction 3 assignees View on GitHub

@carlushuang is already working on this.

Since Aug 28, 2025.

High Priority
Dominant language
Python
Stars
565
Forks
585
Avg merge
3d 4h
Merged PRs (30d)
366

Description

Suggestion Description
  • Adding FP8 KV Cache [Optional] to mla_decode_fwd interface.
  • Per tensor scaling, with precomputed scaling numbers [by quantizer's calibration], first.
  • Per tensor per 128-channel (1x128), with dynamic-quant computed tile-based scaling numbers, second
  • Support FP8 KV Cache transfer in PD for MLA decode
Operating System

No response

GPU

No response

ROCm Component

No response

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.