How does folding constant help improve the trt engine performance
@BowenFu is already working on this.
Since May 7, 2023.
- Dominant language
- C++
- Stars
- 13.4k
- Forks
- 2.4k
- Avg merge
- 5d 3h
- Merged PRs (30d)
- 2
Description
Hi, tensorrt team! I found if I fold constants when exporting the onnx, the trt engine built based on the onnx would be faster when inferring. Could I get an explanation from you?
Besides, in order to refit an engine, we build a mapping relationship between torch model and tensorrt engine without folding constant because folding constant would break the relationship between them (module names will be different). Do you have any approaches to build a relationship between torch model and tensorrt engine which constants were folded, or if not folding constant is necessary, do you have a method to decrease the inferring time?
Thanks!
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.