Logits Processor Support
- Dominant language
- C
- Stars
- 22.3k
- Forks
- 2.1k
- Avg merge
- 1d 3h
- Merged PRs (30d)
- 4
Description
In vllm, there is [Logits Processor](https://docs.vllm.ai/en/latest/design/logits_processors/) API that can let CPU choose resulting token by final logits.
This is a powerful way to constraint tag completeness / reduce tool call mistakes, especially for quantized models.
SGLang was started as a framework for constraining generated tag tokens by context free grammar.
OpenAI's API has a ["lark cfg"](https://developers.openai.com/api/docs/guides/function-calling#lark-cfg) feature for function calling too.
For DS I think exposing a `logits_processor` callback will make the model a lot more playable. Some use case are:
- Perfectly parsable JSON in generated output
- Prevent listing hallucinated file names
- [Control thinking with a grammar](https://www.reddit.com/r/LocalLLaMA/comments/1sx7w55/gbnf_grammar_tweak_for_faster_qwen36_35ba3b_and/)
- Force the answer [cite from prompt](https://github.com/NVIDIA/logits-processor-zoo/blob/main/logits_processor_zoo/vllm/cite_prompt.py)
Contributor guide
Assessment
This issue has not been assessed yet.