microsoft/MInference
View on GitHub[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
- Stars
- 1.2k
- Forks
- 82
- Open beginner issues
- 0
- Indexed issues
- 86
- Avg merge
- 1d 18h
- Merged PRs (30d)
- 1
- Dominant language
- Python
- License
- MIT
- Last GitHub push
- Sep 10, 2026
- Latest indexed
- Sep 19, 2026
- Contributing guide
- No contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- No beginner labels indexed
0 beginner-friendly issues open
Loading issues
No issues to show. Show everything we have indexed