Logits Processor Support
- 主要言語
- C
- スター
- 22.3k
- フォーク
- 2.1k
- 平均マージ
- 1日 3時間
- マージ済み PR(30日)
- 4
説明
In vllm, there is [Logits Processor](https://docs.vllm.ai/en/latest/design/logits_processors/) API that can let CPU choose resulting token by final logits.
This is a powerful way to constraint tag completeness / reduce tool call mistakes, especially for quantized models.
SGLang was started as a framework for constraining generated tag tokens by context free grammar.
OpenAI's API has a ["lark cfg"](https://developers.openai.com/api/docs/guides/function-calling#lark-cfg) feature for function calling too.
For DS I think exposing a `logits_processor` callback will make the model a lot more playable. Some use case are:
- Perfectly parsable JSON in generated output
- Prevent listing hallucinated file names
- [Control thinking with a grammar](https://www.reddit.com/r/LocalLLaMA/comments/1sx7w55/gbnf_grammar_tweak_for_faster_qwen36_35ba3b_and/)
- Force the answer [cite from prompt](https://github.com/NVIDIA/logits-processor-zoo/blob/main/logits_processor_zoo/vllm/cite_prompt.py)
コントリビューションガイド
評価
この issue はまだ評価されていません。