vllm-project / vllm-project/aibrix

Documented example for autoscaling a multinode deployment

Open
#1,799 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
5.1k
Forks
694
Avg merge
1d 19h
Merged PRs (30d)
98

Description

### 🚀 Feature Description and Motivation

I can see in the documentation information for how to use autoscaling and how to do multi-node deployments, but not on both at the same time. The reason I ask for this is that multi-node deployments require scaling of both the head and worker nodes, and I am unsure if this affects how the autoscaler needs to be setup.

The closest to a documented example is in an [issue](https://github.com/vllm-project/aibrix/issues/986).

Does anyone have a working example of a metrics-based autoscaling multi-node deployment?

### Use Case

Having a scalable multi-node deployment would help optimise cost and help handle spikes on demands

### Proposed Solution

_No response_

Contributor guide

Open the contributing guide

Research direction

Start by reading the existing documentation sections on autoscaling and multi-node deployments, then review the related issue #986. Verify how a metrics-based setup scales both head and worker nodes. Done means the documentation contains a working combined example and explains any autoscaler configuration specific to multi-node deployments.

Written by the indexing model from the issue text.

Assessment

Domain
documentation, infrastructure
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.