NVIDIA / NVIDIA/TensorRT

TensorRT 10.3 FP16 inference produces incorrect segmentation results for DINOv3 ViT models while TensorRT 8.6.1 works correctly

Open
#4,723 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Module:Accuracy
Dominant language
C++
Stars
13.4k
Forks
2.4k
Avg merge
5d 3h
Merged PRs (30d)
2

Description

Description

I trained a semantic segmentation model based on DINOv3 ViT backbone using the Lightly_train framework. The model uses pretrained weights from ViT-B16 and ViT-L16 and is fine-tuned for a segmentation task.
However, I encountered inconsistent inference results when converting the model to TensorRT with different TensorRT versions.

Case 1: ViT-B16 backbone

When converting the model to FP16 TensorRT engines:

  • TensorRT 8.6.1

    • Engine builds successfully
    • FP16 inference results are correct
  • TensorRT 10.3

    • Engine also builds successfully
    • But inference results are incorrect
    • The predicted segmentation map contains almost a single class value across the entire image

I verified that:

  • The ONNX model produces correct results
  • The inference code is identical between TensorRT versions
  • CNN-based segmentation models run correctly on TensorRT 10.3 using the same pipeline

Therefore, the issue seems specific to Transformer-based models (ViT / DINOv3).
After inspecting the TensorRT engine behavior, I suspect there may be numerical instability or FP16 overflow in TensorRT 10.3, possibly related to attention or LayerNorm operations.

Case 2: ViT-L16 backbone

For a larger model using ViT-L16 weights, the situation is worse.

When converting to FP16 TensorRT engines:

  • TensorRT 8.6.1
  • TensorRT 10.3
  • TensorRT 10.15

All versions build successfully, but inference results are incorrect.

From preliminary analysis, this may be caused by precision overflow or numerical instability in FP16, since ViT-L16 has a deeper Transformer architecture.

Observations
During conversion, I also noticed a difference in attention handling:

  1. In TensorRT 8.6.1, multi-head attention seems to be fused into optimized kernels
  2. In TensorRT 10.3, attention operations appear to be fully decomposed into MatMul / Transpose / Softmax layers
    This difference might be related to the observed inference errors.

Questions

1、Why does TensorRT 8.6.1 produce correct results while TensorRT 10.3 produces incorrect results for the same ViT-B16 model?

2、Could this be related to:

  • numerical instability in FP16
  • attention decomposition in TensorRT 10.x
  • LayerNorm precision issues?

3、For deeper Transformer models such as ViT-L16, what is the recommended way to build stable FP16 TensorRT engines?

For example:

  • Should certain layers (e.g., LayerNorm or Softmax) be forced to FP32?
  • Are there recommended TensorRT build flags or plugins for Transformer models?

Any guidance would be greatly appreciated.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by comparing the ONNX outputs with FP16 TensorRT inference for the ViT-B16 and ViT-L16 cases across TensorRT 8.6.1, 10.3, and 10.15. Inspect the reported difference between fused multi-head attention and decomposed MatMul/Transpose/Softmax operations, along with possible LayerNorm or FP16 overflow. Done means identifying the numerical cause or documenting a stable build configuration.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.