NVIDIA/Model-Optimizer
View on GitHubA unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
- Stars
- 3.8k
- Forks
- 604
- Open beginner issues
- 0
- Indexed issues
- 93
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 142
- Dominant language
- Python
- License
- Apache-2.0
- Last GitHub push
- Sep 19, 2026
- Latest indexed
- Sep 19, 2026
- Contributing guide
- Contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- No beginner labels indexed
1 beginner-friendly issue open
Loading issues
-
Puzzletron Progress 6/8 (calculating one block scores) takes 10 to 20 times more than in tutorial Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/Model-Optimizer#1667 ·