NVIDIA / NVIDIA/Model-Optimizer

Quantization output using different ONNX Execution Providers

Open
#1,032 1 comment 1 reaction 1 assignee View on GitHub

@ajrasane is already working on this.

Since Mar 16, 2026.

onnx.quantization question triaged
Dominant language
Python
Stars
3.8k
Forks
604
Avg merge
2d 8h
Merged PRs (30d)
142

Description

Hi,

Firstly, thank you for creating this library.

I have 3 questions.

1) What is the difference in an output quantized ONNX model when using different ONNX Execution Providers ?

I train models and convert them to ONNX in Pytorch. I then create TRT engines (timing caches) using ONNXRuntime using the TRT Execution Provider.

So, I am not sure why I would use the TRT Execution Provider for quantization when I will use that provider to create a timing cache after I quantize a model (the result of quantization is a quantized ONNX model).

Another way of asking:

1) What is the difference among these three ?

pth -> onnx -> quantized onnx (CPU) -> timing cache (ONNXRuntime, TRT EP)
pth -> onnx -> quantized onnx (CUDA EP) -> timing cache (ONNXRuntime, TRT EP)
pth -> onnx -> quantized onnx (TRT EP) -> timing cache (ONNXRuntime, TRT EP)

2) Does quantizing using the TRT EP mean that such quantized ONNX models should only/preferably be used on the GPU family the quantization process was done on ?

Hope I managed to explain what I am confused about.

Thanks!

Edit: Another question. 3) Does the library support only the RTX 40 and 50 series ? (I use an RTX 3090, RTX 6000 ADA and some other RTX 30 GPUs)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.