microsoft

microsoft/LLMLingua

View on GitHub

[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.

Stars
6.7k
Forks
428
Open beginner issues
0
Indexed issues
101
Avg merge
2d 4h
Merged PRs (30d)
1
Dominant language
Python
License
MIT
Last GitHub push
Sep 10, 2026
Latest indexed
Sep 19, 2026
Contributing guide
No contributing guide
Code of conduct
Code of conduct
Beginner labels
No beginner labels indexed
101 open issues indexed Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.