NVIDIA / NVIDIA/TensorRT

MHA FP8 Fusion with TensorRT 10.8

Open
#4,516 0 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Module:Performance Module:Runtime
Dominant language
C++
Stars
13.4k
Forks
2.4k
Avg merge
5d 3h
Merged PRs (30d)
2

Description

Description

I tried to reproduce the FP8 MHA fusion with TensorRT 10.8 but from this example but it seems that the MHA is executed in Half precision from the output logs.
Here are the logs from this command:

 trtexec --loadEngine=vit_base_patch8_224_Opset17.engine \
--profilingVerbosity=detailed --dumpLayerInfo --skipInference &> output.log

tensorrt_mha_fp8.log

Is it expected?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the linked TensorRT 10.8 FP8 MHA fusion example, the provided trtexec command, and tensorrt_mha_fp8.log; compare the engine's detailed layer information with the example workflow. Done means establishing whether the Half-precision MHA output is expected and documenting the relevant explanation or next diagnostic step.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.