NVIDIA / NVIDIA/TensorRT

Does tensorrt support dynamic quantization?

Open
#4,698 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Module:Quantization
Dominant language
C++
Stars
13.4k
Forks
2.4k
Avg merge
5d 3h
Merged PRs (30d)
2

Description

Description

I used quantize_dynamic from onnxruntime to quantize onnx model. But it could not convert to tensorrt plan.
error is Non-zero zero point is not supported. Do you know how to fix it?

[02/11/2026-23:33:24] [E] [TRT] ModelImporter.cpp:138: --- Begin node ---
input: "/model/embeddings/tok_embeddings/Gather_output_0_quantized"
input: "model.embeddings.tok_embeddings.weight_scale"
input: "model.embeddings.tok_embeddings.weight_zero_point"
output: "/model/embeddings/tok_embeddings/Gather_output_0"
name: "/model/embeddings/tok_embeddings/Gather_output_0_DequantizeLinear"
op_type: "DequantizeLinear"

[02/11/2026-23:33:24] [E] [TRT] ModelImporter.cpp:139: --- End node ---
[02/11/2026-23:33:24] [E] [TRT] ModelImporter.cpp:141: ERROR: onnxOpImporters.cpp:1584 In function QuantDequantLinearHelper:
[6] Assertion failed: shiftIsAllZeros(zeroPoint): Non-zero zero point is not supported. Please set kENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLAto enable asymmetric quantization if it is on DLA.

import onnx
from onnxruntime.quantization import quantize_dynamic, QuantType

model_fp32 = 'onnx_models/model.onnx'
model_quant = 'onnx_models/model.quant.opt.onnx'
opt = {
"WeightSymmetric": True,
"ActivationSymmetric": True,
}
quantized_model = quantize_dynamic(model_fp32, model_quant, extra_options=opt)

Environment

TensorRT Version:

NVIDIA GPU:

NVIDIA Driver Version:

CUDA Version:

CUDNN Version:

Operating System:

Python Version (if applicable):

Tensorflow Version (if applicable):

PyTorch Version (if applicable):

Baremetal or Container (if so, version):

Relevant Files

Model link:

Steps To Reproduce

Commands or scripts:

Have you tried the latest release?:

Attach the captured .json and .bin files from TensorRT's API Capture tool if you're on an x86_64 Unix system

Can this model run on other frameworks? For example run ONNX model with ONNXRuntime (polygraphy run <model.onnx> --onnxrt):

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the Python quantize_dynamic flow shown in the issue and inspect TensorRT's ONNX importer handling of the DequantizeLinear node and its non-zero zero point. The issue is resolved when dynamic quantization support or its limitations are established, with the required quantization settings or a documented explanation of the failure.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.