NVIDIA/TensorRT-LLM
View on GitHubTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
- Stars
- 14.7k
- Forks
- 2.8k
- Open beginner issues
- 21
- Indexed issues
- 596
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 489
- Dominant language
- Python
- License
- No license data
- Last GitHub push
- Sep 19, 2026
- Latest indexed
- Sep 19, 2026
- Contributing guide
- Contributing guide
- Code of conduct
- Code of conduct
- Beginner labels
- No beginner labels indexed
-
Customized kernels
NVIDIA/TensorRT-LLM#18335 · 2 comments · 1 assignee ·
-
Customized kernels
NVIDIA/TensorRT-LLM#18333 · 1 comment · 1 assignee ·
-
Customized kernels
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#18297 ·
-
Disaggregated serving
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/TensorRT-LLM#18156 ·
-
RFC
NVIDIA/TensorRT-LLM#18153 · 1 assignee ·
-
Disaggregated serving RFC
NVIDIA/TensorRT-LLM#18151 · 1 assignee ·
-
Inference runtime
Difficulty 4/5 3-5 days Newbie friendliness 58/100
NVIDIA/TensorRT-LLM#18115 · 3 comments ·
-
Speculative Decoding
Difficulty 5/5 Over a week Newbie friendliness 25/100
NVIDIA/TensorRT-LLM#18085 ·
-
LLM API
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
NVIDIA/TensorRT-LLM#18084 ·
-
LLM API
Difficulty 3/5 1-2 days Newbie friendliness 66/100
NVIDIA/TensorRT-LLM#18083 · 3 comments ·
-
LLM API
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
NVIDIA/TensorRT-LLM#17977 ·
-
Customized kernels
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#17928 ·
-
Customized kernels
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/TensorRT-LLM#17927 ·
-
Decoding/Sampling
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
NVIDIA/TensorRT-LLM#17916 ·
-
Pytorch
Difficulty 3/5 1-2 days Newbie friendliness 35/100
NVIDIA/TensorRT-LLM#17914 ·
-
bug Pytorch
Difficulty 3/5 1-2 days Newbie friendliness 74/100
NVIDIA/TensorRT-LLM#17913 · 2 comments ·
-
Pytorch
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NVIDIA/TensorRT-LLM#17910 ·
-
Memory
Difficulty 3/5 1-2 days Newbie friendliness 25/100
NVIDIA/TensorRT-LLM#17909 ·
-
KV-Cache Management
Difficulty 3/5 1-2 days Newbie friendliness 72/100
NVIDIA/TensorRT-LLM#17908 · 1 comment ·
-
Model optimization
Difficulty 4/5 3-5 days Newbie friendliness 52/100
NVIDIA/TensorRT-LLM#17900 ·
-
LLM API
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/TensorRT-LLM#17896 ·
-
feature request RFC
NVIDIA/TensorRT-LLM#17778 · 4 assignees ·
-
Infra
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
NVIDIA/TensorRT-LLM#17765 ·
-
Customized kernels
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TensorRT-LLM#17764 ·
-
Disaggregated serving
Difficulty 3/5 1-2 days Newbie friendliness 72/100
NVIDIA/TensorRT-LLM#17763 ·
-
Testing
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
NVIDIA/TensorRT-LLM#17762 ·
-
Infra
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
NVIDIA/TensorRT-LLM#17761 ·
-
Testing
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TensorRT-LLM#17759 ·
-
Testing
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TensorRT-LLM#17757 ·
-
AutoDeploy
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NVIDIA/TensorRT-LLM#17755 ·
-
AutoDeploy
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA/TensorRT-LLM#17753 ·
-
Windows
Difficulty 1/5 Under an hour Newbie friendliness 90/100
NVIDIA/TensorRT-LLM#17743 ·
-
Frontend
Difficulty 4/5 3-5 days Newbie friendliness 52/100
NVIDIA/TensorRT-LLM#17740 ·
-
Pytorch
Difficulty 5/5 Over a week Newbie friendliness 35/100
NVIDIA/TensorRT-LLM#17723 · 1 comment ·
-
Doc
NVIDIA/TensorRT-LLM#17718 · 1 assignee ·
-
Scale-out
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#17714 ·
-
Disaggregated serving
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#17665 · 3 comments ·
-
General perf Performance
NVIDIA/TensorRT-LLM#17625 · 2 assignees ·
-
Customized kernels
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/TensorRT-LLM#17477 ·
-
[Bug]: Interior control-token rejection can split a tool-call prelude and silently drop tool_calls OpenSpeculative Decoding
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/TensorRT-LLM#17437 · 2 comments ·
-
OpenAI API
Difficulty 4/5 3-5 days Newbie friendliness 52/100
NVIDIA/TensorRT-LLM#17436 · 2 comments ·
-
bug Disaggregated serving
Difficulty 5/5 Over a week Newbie friendliness 25/100
NVIDIA/TensorRT-LLM#17429 · 1 comment ·
-
[Bug]: Streaming and non-streaming reasoning parsers produce different content/reasoning splits Open
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/TensorRT-LLM#17296 ·
-
Customized kernels
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/TensorRT-LLM#17126 · 2 comments ·
-
Disaggregated serving Speculative Decoding
Difficulty 4/5 3-5 days Newbie friendliness 52/100
NVIDIA/TensorRT-LLM#17095 · 2 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TensorRT-LLM#17027 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/TensorRT-LLM#17025 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 55/100
NVIDIA/TensorRT-LLM#17024 · 4 comments ·
-
`trtllm-eval --max_output_length` is silently discarded on text tasks whose yaml sets `max_gen_toks` Open
Difficulty 3/5 1-2 days Newbie friendliness 74/100
NVIDIA/TensorRT-LLM#17022 · 2 comments ·
-
Difficulty 3/5 1-2 days Newbie friendliness 74/100
NVIDIA/TensorRT-LLM#17021 ·