NVIDIA/Model-Optimizer
View on GitHubA unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
- Stars
- 3.8k
- Forks
- 604
- Open beginner issues
- 1
- Indexed issues
- 98
- Avg merge
- 2d 6h
- Merged PRs (30d)
- 138
- Dominant language
- Python
- License
- Apache-2.0
- Last GitHub push
- Sep 19, 2026
- Latest indexed
- Sep 19, 2026
- Contributing guide
- Contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- No beginner labels indexed
-
Puzzletron Progress 6/8 (calculating one block scores) takes 10 to 20 times more than in tutorial Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/Model-Optimizer#1667 ·
-
question
NVIDIA/Model-Optimizer#1666 · 6 comments · 1 assignee ·
-
bug documentation triaged
NVIDIA/Model-Optimizer#1637 · 1 comment · 1 assignee ·
-
auto-fixable bug torch.quantization triaged
NVIDIA/Model-Optimizer#1633 · 1 comment · 1 assignee ·
-
bug torch.quantization triaged
NVIDIA/Model-Optimizer#1615 · 2 comments · 1 assignee ·
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/Model-Optimizer#1577 · 1 comment ·
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
NVIDIA/Model-Optimizer#1539 ·
-
question
Difficulty 5/5 Over a week Newbie friendliness 30/100
NVIDIA/Model-Optimizer#1533 ·
-
bug feature onnx.quantization torch.quantization
NVIDIA/Model-Optimizer#1424 · 1 assignee ·
-
feature request
NVIDIA/Model-Optimizer#1396 · 3 comments · 1 assignee ·
-
feature request
NVIDIA/Model-Optimizer#1366 · 1 assignee ·
-
feature request
NVIDIA/Model-Optimizer#1346 · 1 comment · 1 reaction · 2 assignees ·
-
feature request
NVIDIA/Model-Optimizer#1308 · 24 comments · 1 reaction · 2 assignees ·
-
Attention-only LoRA fine-tuning on NVFP4-quantized 122B MoE — community implementation + preprint Openquestion
NVIDIA/Model-Optimizer#1294 · 1 comment · 1 assignee ·
-
question
NVIDIA/Model-Optimizer#1255 · 4 comments · 1 assignee ·
-
feature request
NVIDIA/Model-Optimizer#1237 · 2 comments · 1 assignee ·
-
feature request torch.quantization triaged
NVIDIA/Model-Optimizer#1173 · 1 comment · 1 reaction · 1 assignee ·
-
bug triaged
NVIDIA/Model-Optimizer#1147 · 1 comment · 1 assignee ·
-
[Model] Qwen3TTS Openfeature request onnx.quantization torch.quantization triaged
NVIDIA/Model-Optimizer#1090 · 2 comments · 1 assignee ·
-
Speculative Decoding Opendocumentation question torch.speculative triaged
NVIDIA/Model-Optimizer#1066 · 1 comment · 1 assignee ·
-
bug torch.quantization triaged
NVIDIA/Model-Optimizer#1064 · 1 comment · 1 assignee ·
-
onnx.quantization question triaged
NVIDIA/Model-Optimizer#1032 · 1 comment · 1 reaction · 1 assignee ·
-
bug torch.sparsity
NVIDIA/Model-Optimizer#821 · 1 comment · 1 assignee ·
-
feature request torch.quantization
NVIDIA/Model-Optimizer#797 · 2 comments · 3 reactions · 1 assignee ·
-
bug
NVIDIA/Model-Optimizer#775 · 1 reaction · 1 assignee ·
-
question
NVIDIA/Model-Optimizer#770 · 7 comments · 1 assignee ·
-
investigating model support question torch.quantization
NVIDIA/Model-Optimizer#716 · 5 comments · 1 reaction · 1 assignee ·
-
feature request investigating model support torch.quantization
NVIDIA/Model-Optimizer#647 · 4 comments · 1 assignee ·
-
bug investigating
NVIDIA/Model-Optimizer#628 · 1 comment · 1 assignee ·
-
bug feature investigating onnx.quantization
NVIDIA/Model-Optimizer#621 · 1 comment · 1 assignee ·
-
feature investigating onnx.quantization question
NVIDIA/Model-Optimizer#619 · 1 assignee ·
-
investigating model support question torch.quantization
NVIDIA/Model-Optimizer#604 · 1 comment · 1 assignee ·
-
bug torch.quantization
NVIDIA/Model-Optimizer#569 · 1 comment · 1 assignee ·
-
feature investigating onnx.quantization question
NVIDIA/Model-Optimizer#565 · 1 comment · 1 assignee ·
-
bug export/deploy investigating model support
NVIDIA/Model-Optimizer#491 · 2 comments · 1 assignee ·
-
documentation feature request investigating torch.quantization
NVIDIA/Model-Optimizer#247 · 1 assignee ·
-
feature request investigating model support torch.quantization
NVIDIA/Model-Optimizer#238 · 2 comments · 1 assignee ·
-
bug investigating model support torch.quantization
NVIDIA/Model-Optimizer#226 · 1 comment · 1 assignee ·
-
bug investigating model support torch.quantization
NVIDIA/Model-Optimizer#212 · 1 comment · 1 assignee ·
-
bug investigating model support needs trt/trtllm triage
NVIDIA/Model-Optimizer#174 · 6 comments · 1 assignee ·
-
feature feature request investigating torch.quantization
NVIDIA/Model-Optimizer#139 · 3 comments · 1 assignee ·
-
feature investigating question torch.quantization
NVIDIA/Model-Optimizer#130 · 1 assignee ·
-
feature feature request investigating torch.quantization
NVIDIA/Model-Optimizer#122 · 3 comments · 1 assignee ·
-
bug investigating model support torch.quantization
NVIDIA/Model-Optimizer#104 · 1 assignee ·
-
export/deploy feature feature request investigating
NVIDIA/Model-Optimizer#96 · 5 comments · 1 assignee ·
-
cache_diffusion Openfeature feature request investigating torch.quantization
NVIDIA/Model-Optimizer#95 · 1 assignee ·
-
feature feature request investigating onnx.quantization
NVIDIA/Model-Optimizer#83 · 2 comments · 1 assignee ·
-
bug investigating model support torch.quantization
NVIDIA/Model-Optimizer#72 · 6 comments · 1 assignee ·