vllm-project / vllm-project/aibrix
a cluster with a custom deployment strategy
Open
- Dominant language
- Go
- Stars
- 5.1k
- Forks
- 697
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 104
Description
How to deploy a cluster based on qwen3-14b with tp=2, replica=2, supporting kv cache offloading, and using a relatively new version of vllm? Is there a corresponding yaml demo for reference? Currently, it seems difficult to configure a yaml that suits one's own needs. Are there any relevant guidelines?
Contributor guide
Research direction
The issue names no files, tests, or entry points. Start by identifying how deployment YAML expresses qwen3-14b, tp=2, replica=2, KV-cache offloading, and the requested vLLM version; done means a working example YAML and clear configuration guidelines.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- yaml
- Domain
- ai, infrastructure
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100