Logits Processor Support
- 主要语言
- C
- 星标
- 22.3k
- 派生
- 2.1k
- 平均合并
- 1 天 3 小时
- 30 天内合并 PR
- 4
描述
In vllm, there is [Logits Processor](https://docs.vllm.ai/en/latest/design/logits_processors/) API that can let CPU choose resulting token by final logits.
This is a powerful way to constraint tag completeness / reduce tool call mistakes, especially for quantized models.
SGLang was started as a framework for constraining generated tag tokens by context free grammar.
OpenAI's API has a ["lark cfg"](https://developers.openai.com/api/docs/guides/function-calling#lark-cfg) feature for function calling too.
For DS I think exposing a `logits_processor` callback will make the model a lot more playable. Some use case are:
- Perfectly parsable JSON in generated output
- Prevent listing hallucinated file names
- [Control thinking with a grammar](https://www.reddit.com/r/LocalLLaMA/comments/1sx7w55/gbnf_grammar_tweak_for_faster_qwen36_35ba3b_and/)
- Force the answer [cite from prompt](https://github.com/NVIDIA/logits-processor-zoo/blob/main/logits_processor_zoo/vllm/cite_prompt.py)
贡献指南
评估
这个 Issue 还没有评估数据。