cncf / cncf/mentoring

[CNCF LFX Proposal] Evolve SparkClient into Kubeflow's Unified Data Processing Layer Term 3 (Sep-Nov)

Open
#1,975 21 comments 5 reactions 0 assignees View on GitHub
2026 CNCF Approved Exported lfx mentorship Maintainer/Contribex Approved Mentors Confirmed Proposal Term 3: Sept-Nov Validation Passed
Dominant language
JavaScript
Stars
3.1k
Forks
816
Avg merge
12h 32m
Merged PRs (30d)
8

Description

### CNCF Project

Kubeflow

### Term

2026 Term 3 (Sep-Nov)

### Program Name

Evolve SparkClient into Kubeflow's Unified Data Processing Layer

### Program Description

# Description:

Kubeflow SparkClient (KEP-107) gives Python users a simple way to run Apache Spark on Kubernetes: interactive Spark Connect sessions via `connect()`, batch jobs via `submit_job()`, lifecycle APIs, and a pluggable backend (Kubernetes, future scope - gateway/livy). The foundation lands this term through KEP-107 and GSoC 2026. This program evolves SparkClient into the unified, Pythonic data-processing layer of the Kubeflow SDK, so users go from raw data to trained models through one consistent API, alongside Trainer, Pipelines, Katib, and Model Registry.

The mentee advances four connected workstreams:

1. **Production readiness & observability** — surface metrics, events, and structured logs through the client; add debugging tooling and a richer job/session status model for running Spark at scale.
2. **Kubeflow ecosystem integration** — Integrate with Kubeflow ecosystem and components like Kubeflow Notebook, Kubeflow pipeline, and other. Explore to make SparkClient the SDK's data layer. Flagship deliverable: hand a dataset written by SparkClient to Kubeflow Trainer via a DataFusion-backed cache, with hooks toward Pipelines, Katib, and Model Registry.
4. **Multi-mode execution** — extend beyond interactive/batch to scheduled (`ScheduledSparkApplication`) and streaming jobs behind one consistent Python API. We would like to make it for simple for ML users but making sure advance usecases are also available for data engineers.
5. **Spark ML experience** — Explore Spark MLlib workflows (preprocessing, training, evaluation) through SparkClient, registering results to Model Registry.
6. Writing blogs, examples, and improve IT, unit tests as well as making user AI assistants are helping well writing Spark Client.

# Expected Outcome:

- Observability APIs on SparkClient: job/session metrics, events, and log streaming; an expanded status model; a debugging guide.
- A working Spark → Trainer data handoff through a DataFusion cache, with a runnable end-to-end example.
- Scheduled and streaming execution modes added to SparkClient, with tests and docs.
- Explore Spark MLlib workflow example wired through SparkClient into Model Registry.
- KEP-107 updates, unit and integration tests, and user-facing documentation for each capability and what to add in AI tools/Kubeflow MCP for easy usage for users.

### Technologies

Python, Java, Apache Spark, Kubernetes, distributed data processing, ML workflows; nice-to-have: DataFusion/Arrow, Prometheus/observability tooling

### Skills same as Technologies?

- [x] Yes, the required skills are the same as the technologies listed above.

### Required/Desirable Skills

_No response_

### Mentors

Shekhar Rajak | @shekharrajak | rshekhar.prasad@gmail.com | Shekharrajak
Tariq Hasan | @tariq-hasan| mmtariquehsn@gmail.com | tariq-hasan
Rishabh Singh | @RobuRishabh | roburishabh@outlook.com| RobuRishabh

### Upstream Issue URL

https://github.com/kubeflow/sdk/issues/655

### Application Prerequisites

- [x] Resume
- [x] Cover Letter
- [ ] School Enrollment Verification
- [ ] Participation Permission from school or employer
- [ ] Coding Challenge
- [ ] Custom Prerequisite (fill in details below)

### Coding Challenge URL

_No response_

### Custom Prerequisite Name

_No response_

### Custom Prerequisite Description

_No response_

### Custom Prerequisite — File Upload

- [ ] Yes — completion of this task requires the mentee to submit a file.

---
**LFX program:** [CNCF - Kubeflow: Evolve SparkClient into Kubeflow's Unified Data Processing Layer (2026 Term 3)](https://mentorship.lfx.linuxfoundation.org/project/01d5da81-e5d6-4693-920c-e0e6f4fbc9a8)

Contributor guide

Open the contributing guide

Research direction

The proposal names KEP-107, SparkClient, and upstream Kubeflow SDK issue #655; start by reading those materials and the current SparkClient scope. Done means the agreed observability APIs, Spark-to-Trainer handoff, scheduled and streaming modes, Spark ML example, tests, documentation, and AI-tool integration are delivered.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, kubernetes, prometheus, python
Domain
backend-api-design, data-engineering, distributed-systems, documentation, machine-learning, observability-sre, testing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.