NVIDIA

NVIDIA/Model-Optimizer

View on GitHub

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Stars
3.8k
Forks
604
Open beginner issues
1
Indexed issues
98
Avg merge
2d 6h
Merged PRs (30d)
138
Dominant language
Python
License
Apache-2.0
Last GitHub push
Sep 19, 2026
Latest indexed
Sep 19, 2026
Contributing guide
Contributing guide
Code of conduct
Code of conduct
Beginner labels
No beginner labels indexed
98 open issues indexed Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.