vllm-project / vllm-project/aibrix
Change the batch job management to application master way
- Dominant language
- Go
- Stars
- 5.1k
- Forks
- 697
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 104
Description
### π Feature Description and Motivation
I'd like to borrow the idea from YARN to solve some problems in batch jobs.
- Current: Metadata service orchestrates everything. Worker Job = dumb executor, just sends requests to a given endpoint
- YARN AM: Worker Job is the intelligent coordinator. Metadata Service manage control plan. Worker Job is responsible for spawning job workers. and then execute batch requests. resource negotiation and gpu worker lifecycle could be managed by Job API.
## Key Advantages
1. Decoupled lifecycle: AM starts immediately (CPU is cheap/fast). Resource acquisition happens asynchronously. No need for metadata service to block on "wait for Deployment/StormService ready."
2. AM has local state: The AM can track which requests succeeded/failed, manage retries, adjust concurrency β without constant round-trips to metadata service. Current architecture stores all state in Redis/annotations.
4. Clean preemption: When preempted, the AM saves a checkpoint (which requests completed), releases resources via RM API, and exits. When restarted, it picks up from the checkpoint.
5. Observable: AM exposes its own /status endpoint. Console/monitoring can scrape it directly. No need to reconstruct state from K8s annotations, we probably can use DB to sync status as well.
```
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββββββ
β Aspect β Before β After (YARN AM) β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββ€
β Who creates StormService? β Metadata service β Metadata service (RM), but requested by AM β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββ€
β Who waits for ready? β Metadata service β AM (worker Job) β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββ€
β Who manages progress? β Metadata service via kopf + Redis β AM locally + reports to RM β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββ€
β Who handles preemption? β Metadata service signals worker β RM signals AM, AM checkpoints β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββ€
β Worker Job is... β Dumb request executor β Intelligent job coordinator β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββ€
β Metadata service is... β Orchestrator + executor β Resource Manager + API surface β
βββββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββββββββ
```
### Use Case
better architecture design for batch users.
### Proposed Solution
_No response_
Contributor guide
Research direction
Start by mapping the current metadata service, Worker Job, Redis or Kubernetes-annotation state, and Job API interactions described in the issue. Review how StormService readiness, progress, preemption, and retries are handled today. Done requires an agreed application-master architecture and an implementation plan with defined lifecycle, checkpointing, status, and resource-management behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes, redis
- Domain
- backend-api-design, distributed-systems, infrastructure
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100