vllm-project / vllm-project/aibrix
Documented example for autoscaling a multinode deployment
- Dominant language
- Go
- Stars
- 5.1k
- Forks
- 694
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 98
Description
### 🚀 Feature Description and Motivation
I can see in the documentation information for how to use autoscaling and how to do multi-node deployments, but not on both at the same time. The reason I ask for this is that multi-node deployments require scaling of both the head and worker nodes, and I am unsure if this affects how the autoscaler needs to be setup.
The closest to a documented example is in an [issue](https://github.com/vllm-project/aibrix/issues/986).
Does anyone have a working example of a metrics-based autoscaling multi-node deployment?
### Use Case
Having a scalable multi-node deployment would help optimise cost and help handle spikes on demands
### Proposed Solution
_No response_
Contributor guide
Research direction
Start by reading the existing documentation sections on autoscaling and multi-node deployments, then review the related issue #986. Verify how a metrics-based setup scales both head and worker nodes. Done means the documentation contains a working combined example and explains any autoscaler configuration specific to multi-node deployments.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation, infrastructure
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100