intel/neural-compressor
View on GitHubSOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
- Stars
- 2.7k
- Forks
- 322
- Open beginner issues
- 0
- Indexed issues
- 1
- Avg merge
- 4d 9h
- Merged PRs (30d)
- 25
- Dominant language
- Python
- License
- Apache-2.0
- Last GitHub push
- Sep 16, 2026
- Latest indexed
- Sep 17, 2026
- Contributing guide
- Contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- No beginner labels indexed
-
[Feature] support MXFP W4A8 evaluation on benchmark/vllm-qdq-plugin for accuracy simulation purpose. Open
intel/neural-compressor#2557 · 0 comments · 0 reactions · 1 assignee ·