Lightning-AI / Lightning-AI/lightning-thunder
Support the `trtllm autodeploy` flashinfer KV-Cached attention in Thunder
Open
@kiya00 is already working on this.
Since Aug 26, 2025.
enhancement
- Dominant language
- Python
- Stars
- 1.5k
- Forks
- 121
- PR merge metrics
- No merged PRs in 30d
Description
🚀 Feature
This issue tracks the progress of supporting Flashinfer KV-cached attention that used in TRTLLM autodeploy in Thunder.
We should implement a transformation that replaces the original SDPA with a KV-cached attention.
Additionally, we need to design a way to pass the KV-cache and other required information—such as page numbers—into the computation graph.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.