NVIDIA / NVIDIA/TensorRT

Low ViT Performance Gain on Jetson Thor Using FP8 vs FP16

Open
#4,599 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Module:ONNX Module:Performance Module:Quantization
Dominant language
C++
Stars
13.4k
Forks
2.4k
Avg merge
5d 3h
Merged PRs (30d)
2

Description

Description

Hello,

Looking at the documentation, to enable fp8 operations you need some onnx surgery (inserting Q/DQ at specific locations) to trigger the right MHA (Multi-Head Attention) fusion in conjunction with fp8 precision.

However, the performance improvement is quite low for base ViT model (~20% latency reduction). It is even worse on the EfficientSAM encoder with basically no gain.

By looking at the profiling and layer info from TensorRT the FP8 seems there (even though some tactics are quite cryptic, especially the gmm_mha_v2_#weirdbitstream).

Environment

  • TensorRT Version: 10.13.3
  • NVIDIA GPU: Thor (Jetson DevKit)
  • NVIDIA Driver Version: 580.00
  • CUDA Version: 13

Relevant Files

Steps To Reproduce

Model Optimizer -> commit

ViT-Base FP8 onnx generation:
python3 -m modelopt.onnx.quantization --onnx_path=./vit_base_patch8_224_Opset17.onnx --quantize_mode=fp8 --output_path=./vitb_fp8.onnx

EfficientSAM-S FP8 onnx generation:
python3 -m modelopt.onnx.quantization --onnx_path=./efficientsam_s_encoder.onnx --quantize_mode=fp8 --output_path=./sam_s_fp8.onnx

ViT-Base FP8 engine generation:
trtexec --stronglyTyped --onnx=./vitb_fp8.onnx --saveEngine=./vitb_fp8.engine

ViT-Base FP8 engine generation:
trtexec --stronglyTyped --onnx=./sam_s_fp8.onnx --saveEngine=./sam_s_fp8.engine

TensorRT Layer Info and Profiles

vit_base_patch8_224_Opset17_fp8.json
vit_base_patch8_224_Opset17_fp8.profile.txt
vit_base_patch8_224_Opset17_fp16.json
vit_base_patch8_224_Opset17_fp16.profile.txt
efficientsam_s_encoder_fp8.json
efficientsam_s_encoder_fp8.profile.txt
efficientsam_s_encoder_fp16.json
efficientsam_s_encoder_fp16.profile.txt

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the two Model Optimizer commands and the corresponding trtexec engine builds for ViT-Base and EfficientSAM-S. Compare the linked FP8 and FP16 layer-info JSON files and profile outputs on the specified Jetson Thor environment, focusing on the reported MHA tactics and latency. Done means documenting the cause of the limited FP8 gain and a validated improvement or clear limitation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.