microsoft/LLMLingua
View on GitHub[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
- Stars
- 6.7k
- Forks
- 428
- Open beginner issues
- 0
- Indexed issues
- 101
- Avg merge
- 2d 4h
- Merged PRs (30d)
- 1
- Dominant language
- Python
- License
- MIT
- Last GitHub push
- Sep 10, 2026
- Latest indexed
- Sep 19, 2026
- Contributing guide
- No contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- No beginner labels indexed
0 beginner-friendly issues open
Loading issues
No issues to show. Show everything we have indexed