TRTLLM worker failed when prompt length is large than max_tokens
Open
Nobody has claimed this yet.
bug
Decoding/Sampling
Inference runtime
Scaffolding
- Dominant language
- Python
- Stars
- 14.7k
- Forks
- 2.8k
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 489
Description
TRTLLM worker failed when prompt length is large than max_tokens. We need to handle this issue on scaffolding level.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue identifies the TRTLLM worker and asks for handling at the scaffolding level, but names no files, tests, or expected behavior. Start by locating the worker's prompt and max_tokens handling, reproduce the failure with a prompt longer than max_tokens, and clarify the intended successful behavior before defining completion criteria.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, backend
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100