vllm-project / vllm-project/aibrix

a cluster with a custom deployment strategy

Open
#1,434 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
5.1k
Forks
697
Avg merge
1d 19h
Merged PRs (30d)
104

Description

How to deploy a cluster based on qwen3-14b with tp=2, replica=2, supporting kv cache offloading, and using a relatively new version of vllm? Is there a corresponding yaml demo for reference? Currently, it seems difficult to configure a yaml that suits one's own needs. Are there any relevant guidelines?

Contributor guide

Open the contributing guide

Research direction

The issue names no files, tests, or entry points. Start by identifying how deployment YAML expresses qwen3-14b, tp=2, replica=2, KV-cache offloading, and the requested vLLM version; done means a working example YAML and clear configuration guidelines.

Written by the indexing model from the issue text.

Assessment

Tech stack
yaml
Domain
ai, infrastructure
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.