vllm-project / vllm-project/aibrix

SLO-Driven Resource Management for vLLM

Open
#755 0 comments 0 reactions 1 assignee Claimed by @DwyaneShi View on GitHub
Dominant language
Go
Stars
5.1k
Forks
694
Avg merge
1d 20h
Merged PRs (30d)
98

Description

### 🚀 Feature Description and Motivation

## Background

Different requests have varying input/output lengths, leading to diverse resource requirements. Currently, when a batch of requests gets scheduled together, it is difficult to guarantee individual users' SLA. To address this, we want to design a solution that provides strong SLO guarantees and manages resource tiers and cost models in an SLO-driven manner. This will require engine co-design efforts.

## Challenges
There're few challenges on the concepts that vLLM and external system doesn't have at this moment.

1. Should we use goodput as the primary metric or rely on simpler single-dimension metrics like TTFT or TPOT?
2. Should we define resource classes based on request profiles
3. How can we ensure fair allocation without underutilizing GPU resources?
4. How to map the SLO to Token Pricing model?

I will skip the proposal part and leave this to public discussion at this moment

### Use Case

support multi-tenant use case

### Proposed Solution

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.