serve: VT_SERVER_MAX_NEW_TOKENS does not cap unset or negative max_tokens
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 423
- Forks
- 53
- Avg merge
- 20h 26m
- Merged PRs (30d)
- 310
Description
Problem
The OpenAI serving layer advertises VT_SERVER_MAX_NEW_TOKENS as an operator ceiling, but both chat and completion handlers call request.to_sampling_params() without passing the ceiling as the serving default. Protocol normalization correctly turns non-positive client limits (including Hermes max_tokens=-1) into std::nullopt; the subsequent serving clamp only reduces an already-positive value. The unset case therefore bypasses the configured ceiling and expands to max_model_len - prompt_len in InputProcessor.
Reproduction
Live Gemma-4 agent request on :8010:
- environment:
VT_SERVER_MAX_NEW_TOKENS=4096 - wire request:
max_tokens=-1 - provider request log:
max_tokens=-1 - a stochastic tool-emission failure generated more than 2,250 tokens without EOS before the client moved on; the configured ceiling was never applied at serving admission.
Source chain on current main:
ChatCompletionRequest::to_sampling_params()maps<=0to the optional serving default.OpenAIServingChatandOpenAIServingCompletioncall it with no default.- Chat's local clamp uses
value_or(cap)only for comparison, but does not assign the cap when the optional is unset. - Completion has no equivalent operator-cap application.
Expected behavior
When VT_SERVER_MAX_NEW_TOKENS > 0, it is a hard serving ceiling for all requests:
- unset/non-positive client max -> cap;
- positive client max above cap -> cap;
- positive client max below cap -> unchanged;
- cap
0-> disabled, preserving the protocol's unbounded-to-context behavior.
Apply the policy consistently to /v1/chat/completions and /v1/completions. Keep protocol normalization unchanged.
Acceptance
- RED-first shared-policy unit test covers all four cases above.
- Both chat and completion handlers use the shared policy.
- Existing protocol tests proving non-positive values are UNSET remain unchanged and green.
- Focused OpenAI protocol/serving/API tests pass.
- Full CPU gate status is reported honestly; unrelated baseline failures are not weakened.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start at ChatCompletionRequest::to_sampling_params(), then inspect OpenAIServingChat and OpenAIServingCompletion to trace how the serving default and cap are applied. Add the RED-first shared-policy unit test for unset, non-positive, above-cap, below-cap, and zero-cap cases, then verify both handlers use it while existing protocol tests remain green.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- api, backend-api-design, testing-qa
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Quiet
- Clarity
- Clearly specified
- Newbie friendliness
- 74/100