NVIDIA / NVIDIA/TensorRT-LLM

Example script fails on 5090

Open
#3,347 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug Infra Testing
Dominant language
Python
Stars
14.7k
Forks
2.8k
Avg merge
2d 23h
Merged PRs (30d)
489

Description

System Info
  • x86_64
  • 64GB of DDR5
  • RTX 5090
  • 32GB
    Running cuda-toolkit-12.8, driver 570.124.04, open.

I install using the instruction from here:

sudo apt-get -y install libopenmpi-dev && pip3 install --upgrade pip setuptools && pip3 install tensorrt_llm --extra-index-url https://pypi.nvidia.com

When I try the provided sanity check, I get this error:

python3 example.py 
Authorization required, but no authorization protocol specified

Authorization required, but no authorization protocol specified

Authorization required, but no authorization protocol specified

Authorization required, but no authorization protocol specified

grogu: Problem obtaining CPU time base, detected to be 1302 pico/cycle, adjusted to safe default 500 picos/cycle
grogu.205852Problem obtaining CPU time base, detected to be 1302 pico/cycle, adjusted to safe default 500 picos/cycle
2025-04-07 14:08:08,992 - INFO - flashinfer.jit: Prebuilt kernels not found, using JIT backend
[TensorRT-LLM] TensorRT-LLM version: 0.18.0
/home/chris/Desktop/fun/new/lib/python3.12/site-packages/torch/cuda/__init__.py:235: UserWarning: 
NVIDIA GeForce RTX 5090 with CUDA capability sm_120 is not compatible with the current PyTorch installation.
The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_70 sm_75 sm_80 sm_86 sm_90.
If you want to use the NVIDIA GeForce RTX 5090 GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/

  warnings.warn(
Loading Model: [1/3]    Downloading HF model
Downloaded model to /home/chris/.cache/huggingface/hub/models--TinyLlama--TinyLlama-1.1B-Chat-v1.0/snapshots/fe8a4ea1ffedaf415f4da2f062534de366a451e6
Time: 0.389s
Loading Model: [2/3]    Loading HF model to memory
160it [00:00, 390.14it/s]
/home/chris/Desktop/fun/new/lib/python3.12/site-packages/torch/cuda/__init__.py:235: UserWarning: 
NVIDIA GeForce RTX 5090 with CUDA capability sm_120 is not compatible with the current PyTorch installation.
The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_70 sm_75 sm_80 sm_86 sm_90.
If you want to use the NVIDIA GeForce RTX 5090 GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/

  warnings.warn(
Time: 0.509s
Loading Model: [3/3]    Building TRT-LLM engine
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[04/07/2025-14:08:22] [TRT] [E] IBuilder::buildSerializedNetwork: Error Code 4: Internal Error (Internal error: plugin node LLaMAForCausalLM/transformer/layers/0/attention/wrapper_L567/gpt_attention_L5501/PLUGIN_V2_GPTAttention_0 requires 423138582912 bytes of scratch space, but only 33680457728 is available. Try increasing the workspace size with IBuilderConfig::setMemoryPoolLimit().
)
Traceback (most recent call last):
  File "/home/chris/Desktop/fun/example.py", line 27, in <module>
    main()
  File "/home/chris/Desktop/fun/example.py", line 14, in main
    llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0")
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm.py", line 165, in __init__
    raise e
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm.py", line 160, in __init__
    self._build_model()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm.py", line 407, in _build_model
    self._engine_dir, self._hf_model_dir = model_loader()
                                           ^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1562, in __call__
    return self._build_model(), self._hf_model_dir
           ^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1691, in _build_model
    build_task(self.get_engine_dir())
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1645, in build_task
    model_loader(engine_dir=engine_dir)
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1168, in __call__
    pipeline()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1108, in __call__
    self.step_forward()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1137, in step_forward
    self.step_handlers[self.counter]()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1427, in _build_engine
    self._engine = build(self.model, copied_build_config)
                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/builder.py", line 1287, in build
    engine = None if build_config.dry_run else builder.build_engine(
                                               ^^^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/_common.py", line 206, in decorated
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/builder.py", line 429, in build_engine
    assert engine is not None, 'Engine building failed, please check the error log.'
           ^^^^^^^^^^^^^^^^^^
AssertionError: Engine building failed, please check the error log.
Who can help?

@juney-nvidia
@byshiue

Information
  • The official example scripts
  • My own modified scripts
Tasks
  • An officially supported task in the examples folder (such as GLUE/SQuAD, ...)
  • My own task or dataset (give details below)
Reproduction

sudo apt-get -y install libopenmpi-dev && pip3 install --upgrade pip setuptools && pip3 install tensorrt_llm --extra-index-url https://pypi.nvidia.com

Then past this python code into a script, example:

from tensorrt_llm import LLM, SamplingParams


def main():

    prompts = [
        "Hello, my name is",
        "The president of the United States is",
        "The capital of France is",
        "The future of AI is",
    ]
    sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

    llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0")

    outputs = llm.generate(prompts, sampling_params)

    # Print the outputs.
    for output in outputs:
        prompt = output.prompt
        generated_text = output.outputs[0].text
        print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")


# The entry point of the program need to be protected for spawning processes.
if __name__ == '__main__':
    main()

Expected behavior

Not sure, see NVIDIA tutorial here: https://nvidia.github.io/TensorRT-LLM/quick-start-guide.html

actual behavior

Authorization required, but no authorization protocol specified

Authorization required, but no authorization protocol specified

Authorization required, but no authorization protocol specified

Authorization required, but no authorization protocol specified

grogu: Problem obtaining CPU time base, detected to be 1302 pico/cycle, adjusted to safe default 500 picos/cycle
grogu.205852Problem obtaining CPU time base, detected to be 1302 pico/cycle, adjusted to safe default 500 picos/cycle
2025-04-07 14:08:08,992 - INFO - flashinfer.jit: Prebuilt kernels not found, using JIT backend
[TensorRT-LLM] TensorRT-LLM version: 0.18.0
/home/chris/Desktop/fun/new/lib/python3.12/site-packages/torch/cuda/__init__.py:235: UserWarning: 
NVIDIA GeForce RTX 5090 with CUDA capability sm_120 is not compatible with the current PyTorch installation.
The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_70 sm_75 sm_80 sm_86 sm_90.
If you want to use the NVIDIA GeForce RTX 5090 GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/

  warnings.warn(
Loading Model: [1/3]    Downloading HF model
Downloaded model to /home/chris/.cache/huggingface/hub/models--TinyLlama--TinyLlama-1.1B-Chat-v1.0/snapshots/fe8a4ea1ffedaf415f4da2f062534de366a451e6
Time: 0.389s
Loading Model: [2/3]    Loading HF model to memory
160it [00:00, 390.14it/s]
/home/chris/Desktop/fun/new/lib/python3.12/site-packages/torch/cuda/__init__.py:235: UserWarning: 
NVIDIA GeForce RTX 5090 with CUDA capability sm_120 is not compatible with the current PyTorch installation.
The current PyTorch install supports CUDA capabilities sm_50 sm_60 sm_70 sm_75 sm_80 sm_86 sm_90.
If you want to use the NVIDIA GeForce RTX 5090 GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/

  warnings.warn(
Time: 0.509s
Loading Model: [3/3]    Building TRT-LLM engine
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[TensorRT-LLM][WARNING] Fall back to unfused MHA for data_type = bf16, head_size = 64, head_size_V = 0, attention_mask_type = causal, attention_input_layout = packed_qkv, num_tokens_per_block = 64, alibi = false, attn_logit_softcapping_scale = false in sm_120.
[04/07/2025-14:08:22] [TRT] [E] IBuilder::buildSerializedNetwork: Error Code 4: Internal Error (Internal error: plugin node LLaMAForCausalLM/transformer/layers/0/attention/wrapper_L567/gpt_attention_L5501/PLUGIN_V2_GPTAttention_0 requires 423138582912 bytes of scratch space, but only 33680457728 is available. Try increasing the workspace size with IBuilderConfig::setMemoryPoolLimit().
)
Traceback (most recent call last):
  File "/home/chris/Desktop/fun/example.py", line 27, in <module>
    main()
  File "/home/chris/Desktop/fun/example.py", line 14, in main
    llm = LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0")
          ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm.py", line 165, in __init__
    raise e
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm.py", line 160, in __init__
    self._build_model()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm.py", line 407, in _build_model
    self._engine_dir, self._hf_model_dir = model_loader()
                                           ^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1562, in __call__
    return self._build_model(), self._hf_model_dir
           ^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1691, in _build_model
    build_task(self.get_engine_dir())
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1645, in build_task
    model_loader(engine_dir=engine_dir)
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1168, in __call__
    pipeline()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1108, in __call__
    self.step_forward()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1137, in step_forward
    self.step_handlers[self.counter]()
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/llmapi/llm_utils.py", line 1427, in _build_engine
    self._engine = build(self.model, copied_build_config)
                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/builder.py", line 1287, in build
    engine = None if build_config.dry_run else builder.build_engine(
                                               ^^^^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/_common.py", line 206, in decorated
    return f(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^
  File "/home/chris/Desktop/fun/new/lib/python3.12/site-packages/tensorrt_llm/builder.py", line 429, in build_engine
    assert engine is not None, 'Engine building failed, please check the error log.'
           ^^^^^^^^^^^^^^^^^^
AssertionError: Engine building failed, please check the error log.
additional notes

n/a

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the provided example.py and its LLM(model="TinyLlama/TinyLlama-1.1B-Chat-v1.0") call, then trace the reported failure through tensorrt_llm/llmapi/llm.py and llm_utils.py. Reproduce the sanity check with the listed RTX 5090, PyTorch, and CUDA versions, and determine whether the example can build its engine without the reported scratch-space failure.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
ai
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.