cncf / cncf/sandbox

[Sandbox] KubeRay

Open
#525 3 comments 0 reactions 2 assignees Claimed by @raravena80 View on GitHub
contribution-agreement/unsigned New
Dominant language
JavaScript
Stars
197
Forks
40
Avg merge
10h 30m
Merged PRs (30d)
5

Description

### Project summary

A toolkit and Kubernetes operator that simplifies the deployment, lifecycle management, and autoscaling of Ray applications on Kubernetes

### Project description

KubeRay is a powerful, open source Kubernetes operator designed to run distributed AI, machine learning, and Python workloads powered by [Ray](https://github.com/ray-project/ray) seamlessly on cloud-native infrastructure. As distributed AI workloads (LLM training, batch inference, online serving, and fine-tuning) become a core cloud native construct, KubeRay bridges the gap between Python-first AI frameworks and Kubernetes-native infrastructure.

KubeRay manages the full lifecycle of Ray clusters through three core Custom Resource Definitions (CRDs):

1. **RayCluster**: Manages cluster creation, teardown, autoscaling, and fault tolerance across heterogeneous node pools (CPUs, GPUs, TPUs).
2. **RayJob**: Handles ephemeral job-based execution by dynamically provisioning a RayCluster, running a job to completion, and tearing down resources to optimize compute efficiency.
3. **RayService**: Enables zero-downtime upgrades, high availability, and multi-version traffic routing for Ray Serve applications.

In addition, KubeRay provides a kubectl-ray CLI plugin, an API server, and native integrations with cloud-native queuing and batch systems (such as Kueue, Volcano, and Apache YuniKorn) to ensure multi-tenant fair scheduling on Kubernetes.

### Project vs Reference Architecture/Implementation

- [x] This is a reusable open source project designed for broad adoption, not a reference architecture, reference implementation, or demonstration of patterns

### Org repo URL (provide if all repos under the org are in scope of the application)

https://github.com/ray-project/

### Project repo URL in scope of application

https://github.com/ray-project/kuberay

### Additional repos in scope of the application

N/A

### Website URL

https://github.com/ray-project/kuberay

### Roadmap

https://github.com/ray-project/kuberay/issues/2501

### Roadmap context

https://docs.google.com/document/d/1oz-onT3BTBruVGN0amgD0cfz8ELAiZatZHmjKcnx42Q/edit?tab=t.0#heading=h.8uojskj6olb9

### Contributing guide

https://github.com/ray-project/kuberay/blob/master/CONTRIBUTING.md

### Code of Conduct (CoC)

https://github.com/ray-project/kuberay/blob/master/CODE_OF_CONDUCT.md

### Adopters

N/A

### Maintainers file

https://github.com/ray-project/kuberay/blob/master/docs/community/PEOPLE.md

### Security policy file

https://github.com/ray-project/kuberay/blob/master/SECURITY.md

### Standard or specification?

N/A

### Business product or service to project separation

KubeRay was originally co-developed and maintained by open-source contributors across multiple organizations (including Anyscale, Google, Microsoft, ByteDance, and AWS). While commercial vendors offer managed Ray or AI platform services that utilize KubeRay upstream, KubeRay operates strictly as a vendor-neutral, community-driven Kubernetes operator. Donating KubeRay to the CNCF will ensure long-term governance neutrality, encouraging broader cross-cloud and enterprise contribution without single-vendor control.

### Why CNCF?

1. **Neutral Infrastructure Governance:** As Kubernetes becomes the defacto platform for AI/ML workloads, KubeRay serves as a critical control plane link. CNCF stewardship guarantees vendor-neutral governance and standard IP management (Apache 2.0).
2. **Deep Ecosystem Collaboration:** Bringing KubeRay into the CNCF facilitates tighter co-development with related cloud-native scheduling, observability, and storage projects (e.g., Kueue, Prometheus, Jaeger, Argo Workflows).
3. **Enterprise Confidence:** End-users deploying mission-critical AI/ML pipelines at scale benefit from CNCF's operational maturity, security audits, and community standards.

### Benefit to the landscape

KubeRay fills a key capability gap in the CNCF Machine Learning & AI landscape by providing a standardized, production-grade Kubernetes-native pattern for managing Ray.
Unlike general-purpose job schedulers or batch operators, KubeRay is explicitly tailored to handle the nuances of Ray’s stateful head/worker architecture, actor-based dynamic placement, fine-grained GPU/TPU resource allocation, and zero-downtime model serving. It provides the cloud-native substrate needed to build AI platforms on top of vanilla Kubernetes.

### Cloud native 'fit'

KubeRay is built natively on standard Kubernetes building blocks:

* Implemented as a custom controller using controller-runtime and CustomResourceDefinitions (RayCluster, RayJob, RayService).
* Fits directly into cloud-native GitOps workflows (ArgoCD, Flux) and Helm-based deployments.
* Adheres to Kubernetes operator best practices regarding status subresources, pod templates, autoscaling integration (Cluster Autoscaler / KEDA), and graceful termination.

### Cloud native 'integration'

KubeRay integrates extensively with the CNCF ecosystem:

* **Batch Scheduling & Queuing:** Integrates with CNCF-hosted or affiliated batch schedulers like **Kueue**, **Volcano**, and **Apache YuniKorn** for gang scheduling, multi-tenant quota management, and fair sharing.
* **Observability:** Exposes Native **Prometheus** metrics for Ray cluster states, pod status, and resource usage, along with **Grafana** dashboard templates.
* **Ingress & Networking:** Leverages standard Kubernetes Ingress controllers (NGINX, Envoy Gateway) and Service Meshes (Istio) for exposing Ray Dashboard and Ray Serve endpoints.
* **Tooling:** Implements the kubectl plugin framework via kubectl-ray for direct developer interactions.

### Cloud native overlap

Overlaps partially with Cluster API itself KubeAid CLI is a higher-level orchestrator built on top of CAPI, not a replacement for it; it doesn't compete with CAPI so much as remove the need to hand-write CAPI manifests and glue tooling directly.

### Similar projects

Cluster API Operator, kOps, Kubespray, Rancher/RKE2 provisioning tooling, k0rdent, Sidero/Talos (bare-metal-focused). KubeAid CLI's differentiation is the GitOps-first design and the single-binary, no-long-lived-management-plane pivot model for cloud providers.

### Landscape

https://l.cncf.io/?item=ai-native-infra--orchestration-and-scheduling--kuberay

### Insights

https://insights.linuxfoundation.org/project/kuberay

### Trademark and accounts

- [x] If the project is accepted, I agree to donate all project trademarks and accounts to the CNCF

### IP policy

- [x] If the project is accepted, I agree the project will follow the CNCF IP Policy

### Will the project require a license exception?

N/A project uses Apache 2.0.

### Does the project utilize any dependencies with licenses not on the allowlist or an approved exception?

N/A project uses Apache 2.0.

### Project "Domain Technical Review"

N/A

### Application contact email(s)

rliaw@anyscale.com, rkn@anyscale.com, rueian@anyscale.com, kaihsun2019@gmail.com, andrewsy@google.com, eoakes@anyscale.com

### Contributing or sponsoring entity signatory information

| Name | Address | Type (e.g., Delaware corporation) | Signatory name and title | Email address |
|-----------|-----------|-----------|-----------|-----------|
| Anyscale Inc | 600 Harrison St, San Francisco CA 94107 | Delaware corporation | Richard Liaw, Product Manager | rliaw@anyscale.com |
| Anyscale Inc | 600 Harrison St, San Francisco CA 94107 | Delaware corporation | Robert Nishihara, Co-founder | rkn@anyscale.com |

### CNCF contacts

* Chris Aniszczyk

### Additional information

* Vote: https://github.com/ray-project/ray/issues/65900

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.