NVIDIA / NVIDIA/TensorRT

Incorrect results of TensorRT 10.16.1 when running with `--best` optimized resnet152 on GPU RTX PRO 6000 Blackwell

Open
#4,780 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Module:Accuracy
Dominant language
C++
Stars
13.4k
Forks
2.4k
Avg merge
5d 3h
Merged PRs (30d)
2

Description

Description

When using torchvision built-in resnet152 with DEFAULT weights (i.e., ResNet152_Weights.IMAGENET1K_V2), export it to ONNX with torch.onnx.export, and convert to tensorrt engine with trtexec --best and run it, the inference accuracy drops to near zero.

This doesn't happen when the weights are changed to ResNet152_Weights.IMAGENET1K_V1. I've also changed the batch size, the image preprocessing pipeline, yet the issue persists.

Environment

TensorRT Version: 10.16.1.11

NVIDIA GPU: NVIDIA RTX PRO 6000 Blackwell

NVIDIA Driver Version: 580.142

CUDA Version: 13.2

CUDNN Version: 9.10.2

Operating System: Ubuntu 24.04 LTS

Python Version (if applicable): 3.12.3

Tensorflow Version (if applicable): N/A

PyTorch Version (if applicable): 2.12.0a0+0291f960b6.nv26.4.48445190

Baremetal or Container (if so, version): nvcr.io/nvidia/pytorch:26.04-py3 docker image.

Relevant Files

Steps To Reproduce

Download the attached file: repro_tensorrt_best_resnet152.py.

Commands or scripts:

docker run --rm -it --gpus all \
  --ipc=host --ulimit memlock=-1 --ulimit stack=67108864 \
  -v $PWD:/workspace \
  -v .cache:/root/.cache \
  nvcr.io/nvidia/pytorch:26.04-py3 \
  bash -lc "pip3 install git+https://github.com/modestyachts/ImageNetV2_pytorch && python3 repro_tensorrt_best_resnet152.py"

Note the ~0% accuracy issue:

imagenet1k_v2: ResNet152_Weights.IMAGENET1K_V2
pytorch: top-1=71.71%, top-5=89.98%
trt_fp16: top-1=71.71%, top-5=89.99%
trt_best: top-1=0.10%, top-5=0.48%

imagenet1k_v1: ResNet152_Weights.IMAGENET1K_V1
pytorch: top-1=66.96%, top-5=87.22%
trt_fp16: top-1=66.95%, top-5=87.25%
trt_best: top-1=66.00%, top-5=86.64%

Both weights optimized with trtexec --best have the following warnings, but only the V2 weights will have near-zero accuracy:

[W] [TRT] Dequantize 1 [SCALE] has invalid precision Int8, ignored.

I'm suspecting this may be due to not providing a calibration file, but I'm unsure why it differs so much in V1 and V2 weights.

Have you tried the latest release?: Yes

Attach the captured .json and .bin files from TensorRT's API Capture tool if you're on an x86_64 Unix system

Can this model run on other frameworks? For example run ONNX model with ONNXRuntime (polygraphy run <model.onnx> --onnxrt): Yes

# polygraphy run resnet152.onnx --onnxrt
[W] 'colored' module is not installed, will not use colors when logging. To enable colors, please install the 'colored' module: python3 -m pip install colored
[I] RUNNING | Command: /usr/local/bin/polygraphy run resnet152.onnx --onnxrt
[I] onnxrt-runner-N0-05/17/26-17:42:11  | Activating and starting inference
[I] Creating ONNX-Runtime Inference Session with providers: ['CPUExecutionProvider']
[I] onnxrt-runner-N0-05/17/26-17:42:11 
    ---- Inference Input(s) ----
    {input [dtype=float32, shape=(50, 3, 224, 224)]}
[I] onnxrt-runner-N0-05/17/26-17:42:11 
    ---- Inference Output(s) ----
    {logits [dtype=float32, shape=(50, 1000)]}
[I] onnxrt-runner-N0-05/17/26-17:42:11  | Completed 1 iteration(s) in 5762 ms | Average inference time: 5762 ms.
[I] PASSED | Runtime: 7.970s | Command: /usr/local/bin/polygraphy run resnet152.onnx --onnxrt

Disclaimer: The minimal repro code is generated by Codex to keep the code minimal based on a much lengthier codebase.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running repro_tensorrt_best_resnet152.py in the provided nvcr.io/nvidia/pytorch:26.04-py3 container and compare the V1 and V2 results with trtexec --best and FP16. Review the attached trtexec logs, especially the invalid Int8 precision warning, and compare the TensorRT path with the ONNX Runtime result. Done means identifying and correcting the cause of near-zero V2 accuracy without regressing V1.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, python
Domain
backend, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.