NVIDIA / NVIDIA/TensorRT

out of memory failure of TensorRT 10.5 when running flux dit on GPU L40S

Open
#4,214 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Module:Demo triaged
Dominant language
C++
Stars
13.4k
Forks
2.4k
Avg merge
5d 3h
Merged PRs (30d)
2

Description

Description

I tried to convert the Flux Dit model on L40S with TensorRT10.5, and found that the peak gpu memory exceeded 46068MiB, but 23597MiB gpu memory was occupied during inference. Is this normal? If normal, what measures can be taken to reduce the gpu memory usage during model conversion so that Flux TensorRT inference can be run normally in the L40S

[10/17/2024-11:07:02] [I] [TRT] [MemUsageStats] Peak memory usage of TRT CPU/GPU memory allocators: CPU 22681 MiB, GPU 49917 MiB

Environment

TensorRT Version: 10.5

**NVIDIA GPU **: L40S

NVIDIA Driver Version: 535.129.03

CUDA Version: 12.2

CUDNN Version:

Operating System:

Python Version (if applicable):

Tensorflow Version (if applicable):

PyTorch Version (if applicable):

Baremetal or Container (if so, version):

Relevant Files

Model link:

Steps To Reproduce

Commands or scripts:

Have you tried the latest release?:

Can this model run on other frameworks? For example run ONNX model with ONNXRuntime (polygraphy run <model.onnx> --onnxrt):

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the TensorRT 10.5 memory statistics and the reported L40S, driver, and CUDA environment. Reproduce Flux DiT conversion if the model and commands can be obtained; done means determining whether the peak allocation is expected and identifying a supported way to reduce conversion memory or documenting the limitation.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.