AnswerDotAI / AnswerDotAI/cold-compress
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Open
- Dominant language
- Python
- Stars
- 153
- Forks
- 16
- PR merge metrics
- No merged PRs in 30d
Description
Implement this [paper](https://arxiv.org/abs/2407.02490).
Similar to `class KVCacheFastGen` in that it involves a profiling step.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.