NVIDIA / NVIDIA/TensorRT

Different numbers of outputs result in inconsistent running results (original: 不同个数的输出导致运行结果不一致)

Open
#4,284 4 comments 0 reactions 1 assignee View on GitHub

@yuanyao-nv is already working on this.

Since Feb 11, 2025.

Module:ONNX triaged
Dominant language
C++
Stars
13.4k
Forks
2.4k
Avg merge
5d 3h
Merged PRs (30d)
2

Description

Description

Auto-translated

(Part of the linear attention code)

hidden_states = query_KV / query_Z
return hidden_states, query_KV, query_Z

and

hidden_states = query_KV / query_Z
return hidden_states

When converting to TensorRT from ONNX, the results are different between the two above methods. The first is correct, while the second results in NaN; it seems that outputting the intermediate states affects the correctness of the results. How can this issue be resolved?

Further testing shows: This issue occurs with TensorRT 10.7, but TensorRT 10.6 works correctly.

Original

(部分linear attetion代码)
hidden_states = query_KV / query_Z
return hidden_states, query_KV, query_Z

hidden_states = query_KV / query_Z
return hidden_states

上面两者方式onnx转tensorrt时,两者结果不一样,前者是正确的,后者会出现nan;
这相当于输出了中间状态会导致结果的正确性,该怎么解决这种问题哇?

后面测试:tensorrt 10.7会出现这种问题,10.6是正确的

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.