microsoft

microsoft/MInference

View on GitHub

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

Stars
1.2k
Forks
82
Open beginner issues
0
Indexed issues
86
Avg merge
1d 18h
Merged PRs (30d)
1
Dominant language
Python
License
MIT
Last GitHub push
Sep 10, 2026
Latest indexed
Sep 19, 2026
Contributing guide
No contributing guide
Code of conduct
Code of conduct
Beginner labels
No beginner labels indexed
86 issues indexed so far Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.