[CNCF LFX Proposal] Evolve SparkClient into Kubeflow's Unified Data Processing Layer Term 3 (Sep-Nov)
- Dominant language
- JavaScript
- Stars
- 3.1k
- Forks
- 816
- Avg merge
- 12h 32m
- Merged PRs (30d)
- 8
Description
### CNCF Project
Kubeflow
### Term
2026 Term 3 (Sep-Nov)
### Program Name
Evolve SparkClient into Kubeflow's Unified Data Processing Layer
### Program Description
# Description:
Kubeflow SparkClient (KEP-107) gives Python users a simple way to run Apache Spark on Kubernetes: interactive Spark Connect sessions via `connect()`, batch jobs via `submit_job()`, lifecycle APIs, and a pluggable backend (Kubernetes, future scope - gateway/livy). The foundation lands this term through KEP-107 and GSoC 2026. This program evolves SparkClient into the unified, Pythonic data-processing layer of the Kubeflow SDK, so users go from raw data to trained models through one consistent API, alongside Trainer, Pipelines, Katib, and Model Registry.
The mentee advances four connected workstreams:
1. **Production readiness & observability** — surface metrics, events, and structured logs through the client; add debugging tooling and a richer job/session status model for running Spark at scale.
2. **Kubeflow ecosystem integration** — Integrate with Kubeflow ecosystem and components like Kubeflow Notebook, Kubeflow pipeline, and other. Explore to make SparkClient the SDK's data layer. Flagship deliverable: hand a dataset written by SparkClient to Kubeflow Trainer via a DataFusion-backed cache, with hooks toward Pipelines, Katib, and Model Registry.
4. **Multi-mode execution** — extend beyond interactive/batch to scheduled (`ScheduledSparkApplication`) and streaming jobs behind one consistent Python API. We would like to make it for simple for ML users but making sure advance usecases are also available for data engineers.
5. **Spark ML experience** — Explore Spark MLlib workflows (preprocessing, training, evaluation) through SparkClient, registering results to Model Registry.
6. Writing blogs, examples, and improve IT, unit tests as well as making user AI assistants are helping well writing Spark Client.
# Expected Outcome:
- Observability APIs on SparkClient: job/session metrics, events, and log streaming; an expanded status model; a debugging guide.
- A working Spark → Trainer data handoff through a DataFusion cache, with a runnable end-to-end example.
- Scheduled and streaming execution modes added to SparkClient, with tests and docs.
- Explore Spark MLlib workflow example wired through SparkClient into Model Registry.
- KEP-107 updates, unit and integration tests, and user-facing documentation for each capability and what to add in AI tools/Kubeflow MCP for easy usage for users.
### Technologies
Python, Java, Apache Spark, Kubernetes, distributed data processing, ML workflows; nice-to-have: DataFusion/Arrow, Prometheus/observability tooling
### Skills same as Technologies?
- [x] Yes, the required skills are the same as the technologies listed above.
### Required/Desirable Skills
_No response_
### Mentors
Shekhar Rajak | @shekharrajak | rshekhar.prasad@gmail.com | Shekharrajak
Tariq Hasan | @tariq-hasan| mmtariquehsn@gmail.com | tariq-hasan
Rishabh Singh | @RobuRishabh | roburishabh@outlook.com| RobuRishabh
### Upstream Issue URL
https://github.com/kubeflow/sdk/issues/655
### Application Prerequisites
- [x] Resume
- [x] Cover Letter
- [ ] School Enrollment Verification
- [ ] Participation Permission from school or employer
- [ ] Coding Challenge
- [ ] Custom Prerequisite (fill in details below)
### Coding Challenge URL
_No response_
### Custom Prerequisite Name
_No response_
### Custom Prerequisite Description
_No response_
### Custom Prerequisite — File Upload
- [ ] Yes — completion of this task requires the mentee to submit a file.
---
**LFX program:** [CNCF - Kubeflow: Evolve SparkClient into Kubeflow's Unified Data Processing Layer (2026 Term 3)](https://mentorship.lfx.linuxfoundation.org/project/01d5da81-e5d6-4693-920c-e0e6f4fbc9a8)
Contributor guide
Research direction
The proposal names KEP-107, SparkClient, and upstream Kubeflow SDK issue #655; start by reading those materials and the current SparkClient scope. Done means the agreed observability APIs, Spark-to-Trainer handoff, scheduled and streaming modes, Spark ML example, tests, documentation, and AI-tool integration are delivered.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, kubernetes, prometheus, python
- Domain
- backend-api-design, data-engineering, distributed-systems, documentation, machine-learning, observability-sre, testing
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100