NVIDIA/TensorRT-LLM
View on GitHubTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
- Stars
- 14.7k
- Forks
- 2.8k
- Open beginner issues
- 21
- Indexed issues
- 596
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 489
- Dominant language
- Python
- License
- No license data
- Last GitHub push
- Sep 19, 2026
- Latest indexed
- Sep 19, 2026
- Contributing guide
- Contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- No beginner labels indexed
-
[Bug] Attention-DP fill-gate fail-fast can skip the model-parallel status gather on a post-fill rank OpenDisaggregated serving
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#19435 ·
-
bug Customized kernels
NVIDIA/TensorRT-LLM#19364 · 1 assignee ·
-
bug Customized kernels
NVIDIA/TensorRT-LLM#19362 · 1 assignee ·
-
Difficulty 4/5 3-5 days Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#19337 · 2 comments ·
-
Scaffolding
Difficulty 4/5 3-5 days Newbie friendliness 25/100
NVIDIA/TensorRT-LLM#19329 ·
-
Scaffolding
Difficulty 5/5 Over a week Newbie friendliness 35/100
NVIDIA/TensorRT-LLM#19319 · 1 comment ·
-
RFC
NVIDIA/TensorRT-LLM#19301 · 3 assignees ·
-
KV-Cache Management
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
NVIDIA/TensorRT-LLM#19281 ·
-
new model
Difficulty 5/5 Over a week Newbie friendliness 30/100
NVIDIA/TensorRT-LLM#19278 ·
-
KV-Cache Management
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/TensorRT-LLM#19242 ·
-
Customized kernels
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#19240 ·
-
Disaggregated serving Frontend Investigating LLM API RFC triaged
NVIDIA/TensorRT-LLM#19232 · 1 comment · 1 reaction · 4 assignees ·
-
LLM API
Difficulty 2/5 1-3 hours Newbie friendliness 52/100
NVIDIA/TensorRT-LLM#19229 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 35/100
NVIDIA/TensorRT-LLM#19227 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 62/100
NVIDIA/TensorRT-LLM#19225 ·
-
[Bug]: standalone Responses API input, streaming, and tool continuation failures with Agentic API OpenLLM API
Difficulty 5/5 Over a week Newbie friendliness 25/100
NVIDIA/TensorRT-LLM#19223 ·
-
Pytorch
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#19193 ·
-
Infra
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#19173 ·
-
bug Inference runtime Triton backend
Difficulty 5/5 Over a week Newbie friendliness 35/100
NVIDIA/TensorRT-LLM#19172 ·
-
Infra
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#19067 ·
-
Customized kernels
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#19063 · 3 comments ·
-
Pytorch
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/TensorRT-LLM#19059 · 3 comments ·
-
feature request Infra
NVIDIA/TensorRT-LLM#19033 · 1 assignee ·
-
[Bug]: Multimodal metadata failures terminate the PyTorch executor instead of failing the request OpenMultimodal Pytorch
Difficulty 3/5 1-2 days Newbie friendliness 72/100
NVIDIA/TensorRT-LLM#18971 · 1 comment ·
-
Speculative Decoding
Difficulty 4/5 3-5 days Newbie friendliness 66/100
NVIDIA/TensorRT-LLM#18892 · 1 comment ·
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/TensorRT-LLM#18891 · 1 comment ·
-
Inference runtime
Difficulty 3/5 1-2 days Newbie friendliness 78/100
NVIDIA/TensorRT-LLM#18848 · 1 comment ·
-
[Bug] Worker CPU affinity is applied process-wide in shared-process deployments and never restored OpenInference runtime
Difficulty 5/5 Over a week Newbie friendliness 35/100
NVIDIA/TensorRT-LLM#18847 ·
-
Inference runtime
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/TensorRT-LLM#18816 ·
-
Doc
Difficulty 3/5 1-2 days Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#18777 ·
-
Disaggregated serving Pytorch
Difficulty 1/5 Under an hour Newbie friendliness 90/100
NVIDIA/TensorRT-LLM#18776 ·
-
bug Customized kernels Windows
Difficulty 3/5 1-2 days Newbie friendliness 75/100
NVIDIA/TensorRT-LLM#18775 · 1 comment ·
-
Inference runtime
Difficulty 3/5 1-2 days Newbie friendliness 78/100
NVIDIA/TensorRT-LLM#18759 ·
-
bug VisualGen
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/TensorRT-LLM#18687 ·
-
Low Precision
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#18664 ·
-
Inference runtime Pytorch
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#18663 · 2 comments ·
-
Speculative Decoding
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#18662 ·
-
KV-Cache Management Pytorch
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#18661 · 1 comment ·
-
KV-Cache Management
NVIDIA/TensorRT-LLM#18660 · 1 assignee ·
-
Customized kernels Model optimization
Difficulty 4/5 3-5 days Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#18659 ·
-
Pytorch
Difficulty 3/5 1-2 days Newbie friendliness 68/100
NVIDIA/TensorRT-LLM#18658 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 35/100
NVIDIA/TensorRT-LLM#18616 ·
-
bug Pytorch
Difficulty 4/5 3-5 days Newbie friendliness 25/100
NVIDIA/TensorRT-LLM#18609 · 2 comments ·
-
NVIDIA/TensorRT-LLM#18502 · 1 assignee ·
-
Disaggregated serving Testing
Difficulty 5/5 Over a week Newbie friendliness 25/100
NVIDIA/TensorRT-LLM#18450 ·
-
bug KV-Cache Management Speculative Decoding
Difficulty 3/5 1-2 days Newbie friendliness 76/100
NVIDIA/TensorRT-LLM#18449 · 2 comments ·
-
KV-Cache Management Lora/P-tuning Pytorch
Difficulty 3/5 1-2 days Newbie friendliness 72/100
NVIDIA/TensorRT-LLM#18407 ·
-
Doc
Difficulty 4/5 3-5 days Newbie friendliness 42/100
NVIDIA/TensorRT-LLM#18380 · 1 comment ·
-
Model optimization
NVIDIA/TensorRT-LLM#18377 · 2 comments · 1 assignee ·
-
Customized kernels
NVIDIA/TensorRT-LLM#18376 · 1 assignee ·