NVIDIA

NVIDIA/Model-Optimizer

View on GitHub

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Stars
3.8k
Forks
604
Open beginner issues
0
Indexed issues
93
Avg merge
2d 8h
Merged PRs (30d)
142
Dominant language
Python
License
Apache-2.0
Last GitHub push
Sep 19, 2026
Latest indexed
Sep 19, 2026
Contributing guide
Contributing guide
Code of conduct
Code of conduct
Beginner labels
No beginner labels indexed
1 beginner-friendly issue open Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.