How to Perform Static Quantization Directly on an ONNX Model Using Intel® Neural Compressor?
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Documentation
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- python, pytorch
- Domain
- documentation, machine-learning
Research direction
Start with the linked quantize_with_inc.ipynb sample and inspect how it saves the quantized PyTorch model. Verify whether the requested ONNX input/output path is supported and whether an export example belongs in this sample. Done means a clear, reproducible answer or example covering direct ONNX quantization and the alternative export path.
Written by the indexing model from the issue text.
Description
Hello,
I'm using Intel® Neural Compressor (INC) to perform static quantization on my custom PyTorch model. I followed this script which demonstrates how to apply static quantization using INC on a PyTorch model.
My goal is to obtain the final quantized model in ONNX format. However, after quantization, saving the q_model results in a .pt file (PyTorch format). I also found that exporting quantized PyTorch models to ONNX is problematic due to limited support and compatibility issues, especially with static quantization.
My Question:
Is there a way to perform static quantization directly on an ONNX model using Intel® Neural Compressor to produce a quantized ONNX model as the output?
Alternatively, is there a specific method to export the statically quantized PyTorch model to ONNX format while addressing the compatibility issues?
Any guidance or examples on how to achieve this would be greatly appreciated.
Thank you!
- Dominant language
- C++
- Stars
- 1.2k
- Forks
- 745
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from oneapi-src/oneAPI-samples
-
bug
Difficulty 3/5 1-2 days Newbie friendliness 45/100
oneapi-src/oneAPI-samples#2764 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
oneapi-src/oneAPI-samples#2748 ·
-
Difficulty 5/5 Over a week Newbie friendliness 10/100
oneapi-src/oneAPI-samples#2746 ·
-
Lenovo Thinkpad T440p Ethernet issue warning sign show install multiple times but same issue face Openquestion
Difficulty 4/5 3-5 days Newbie friendliness 15/100
oneapi-src/oneAPI-samples#2739 ·
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 35/100
oneapi-src/oneAPI-samples#2705 ·
All issues in oneapi-src/oneAPI-samples
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
-
Sensor initialization takes very long when `--initial-sim-time` is set to current UNIX timestamp Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
gazebosim/gz-sensors#662 · 1 comment ·
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
comp-datalake
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
ClickHouse/ClickHouse#121222 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
LadybirdBrowser/ladybird#12123 ·