[Bug] Users need to add `cuda_graph_max_batch_size=0` to avoid crash when config from extra-llm-api-config.yml
Open
@Superjomn is already working on this.
Since May 30, 2025.
bug
Customized kernels
LLM API
- Dominant language
- Python
- Stars
- 14.7k
- Forks
- 2.8k
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 489
Description
System Info
See discussions here:
- https://github.com/NVIDIA/TensorRT-LLM/pull/4603#discussion_r2112560543
- https://github.com/NVIDIA/TensorRT-LLM/pull/4603#discussion_r2116676593
Who can help?
No response
Information
- The official example scripts
- My own modified scripts
Tasks
- An officially supported task in the
examplesfolder (such as GLUE/SQuAD, ...) - My own task or dataset (give details below)
Reproduction
It would crash if set cuda_graph_batch_size but not cuda_graph_max_batch_size
Expected behavior
Improve backward compatibility of extra-llm-api-config.yml
actual behavior
Raise assertion error
additional notes
NA
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.