NVIDIA

NVIDIA/Model-Optimizer

View on GitHub

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Stars
3.8k
Forks
604
Open beginner issues
0
Indexed issues
93
Avg merge
2d 8h
Merged PRs (30d)
142
Dominant language
Python
License
Apache-2.0
Last GitHub push
Sep 19, 2026
Latest indexed
Sep 19, 2026
Contributing guide
Contributing guide
Code of conduct
Code of conduct
Beginner labels
No beginner labels indexed
93 issues indexed so far Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.