[LLM Runner] Wire grammar-constrained decoding end-to-end
Open
@kirklandsign is already working on this.
Since May 11, 2026.
module: llm
triaged
- Dominant language
- Python
- Stars
- 5k
- Forks
- 1.2k
- Avg merge
- 2d 10h
- Merged PRs (30d)
- 581
Description
When GenerationConfig.grammar is set, create GrammarLogitProcessor, inject into TextTokenGenerator, call accept_token() after sampling. E2E test with JSON schema. Target: <100us mask overhead. Depends on: GenerationConfig fields, GrammarLogitProcessor.
cc @larryliu0820 @mergennachin @cccclai @helunwencser @jackzhxng
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.