mudler / mudler/vllm.cpp

vllm-bench cannot express the KV-pool byte budget: --kv-cache-memory / --gpu-memory-utilization unreachable from the bench

Open
#385 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
423
Forks
53
Avg merge
20h 26m
Merged PRs (30d)
310

Description

The gap

Of the three entrypoints, vllm-bench is the only one that cannot express the KV-pool byte budget.

flag vllm-server vllm-cli vllm-bench
--num-blocks yes yes yes
--kv-cache-memory yes yes no
--gpu-memory-utilization yes yes no

server_main.cpp:289-290 and examples/cli/main.cpp:113-114 both parse --kv-cache-memory into EngineParams::kv_cache_memory_bytes. examples/bench/main.cpp parses --num-blocks (:91-92) and neither of the other two.

To be precise about what this is and is not: vllm-bench rejects unknown arguments loudly (main.cpp:99-103: "vllm-bench: unknown argument", usage, exit_code = 2), so nothing is silently ignored. This is a missing feature, not a silent failure.

Why it is more than cosmetic

bench_core.h:574-576 gives the real-checkpoint path its own sizing heuristic:

params.num_blocks = cfg.num_blocks > 0
                        ? cfg.num_blocks
                        : std::max(cfg.concurrency * seq_blocks * 2, 256);

num_blocks is therefore always > 0 from the bench, so ResolveNumBlocks always takes knob 1 and the byte-budget and utilization knobs are unreachable from this binary by construction -- not merely unparsed. Two consequences:

  1. A benchmark cannot be run under the same KV-pool configuration as the server it is meant to characterize. On a memory-constrained or shared device that is exactly the configuration you want to measure.
  2. The bench's heuristic scales with concurrency and sequence length, so its pool size moves with the benchmark parameters. Comparing a c=8 run against a c=1 run silently varies the KV pool between the two arms, which is a confound in a tool whose purpose is A/B measurement.

Suggested shape

Parse --kv-cache-memory (and --gpu-memory-utilization, for symmetry) in examples/bench/main.cpp and thread them into the EngineParams built in bench_core.h, leaving the current heuristic as the fallback when none of the three knobs is set. That keeps every existing invocation byte-identical and makes the precedence in model_loader.h:69-83 reachable from all three entrypoints.

Note this interacts with #357: while the KVBytesPerBlock layer-count defect is open, the byte-budget knob mis-sizes wherever it is reachable. The order probably wants to be #357 first, then this -- otherwise the new flag inherits a broken divisor. Flagging the dependency rather than assuming your sequencing.

Scope

Read from the tree at 6dbedf9f. I have not built or measured anything for this, and I have not hit a failure caused by it -- I found it while ruling out a flag-plumbing hypothesis for #357 (which turned out to be a real defect elsewhere; the server's plumbing is correct). Happy to implement if you want it.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with argument parsing in examples/bench/main.cpp, then follow the EngineParams construction in bench_core.h and precedence in model_loader.h:69-83. Compare the existing handling in server_main.cpp:289-290 and examples/cli/main.cpp:113-114. Done means vllm-bench accepts both KV-pool options, preserves the current heuristic when no knob is supplied, and makes the documented precedence reachable.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
cli, tooling
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.