apache / apache/datafusion-ballista

[EPIC] Improve job data cleanup logic

Open
#1,316 11 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Rust
Stars
2.1k
Forks
320
Avg merge
1d 22h
Merged PRs (30d)
66

Description

**Is your feature request related to a problem or challenge? Please describe what you are trying to do.**

There are several improvement points in the current job-data deletion flow:

1. *Duplication*
The executor has `clean_all_shuffle_data` alongside other ad-hoc removal logic. These overlap in functionality, making the code harder to maintain and reason about.

2. *Push-based broadcast*
When the scheduler initiates cleanup, it currently notifies all executors. This is inefficient because only a subset of executors actually hold the job’s data.

3. *Per-job deletion tasks*
In `clean_up_successful_job` / `clean_up_failed_job`, the scheduler spawns a separate delayed task (`sleep`) for each job and calls `state.remove_job(job_id)` individually. This results in many small tasks and RPCs, which could be batched more efficiently.

**Describe the solution you'd like**
Unify cleanup behind a single, testable “deletion facility”:

1. *Deduplicate* logic with `clean_all_shuffle_data`; extract/keep a shared async remover (e.g., `remove_job_dir`) with safety checks.
2. *Targeted push*: notify only executors that actually hold the job’s data (no broadcast).
3. *Batching*: we already dispatch periodically; change each tick to send one batched `remove_jobs(Vec)` for all pending IDs rather than spawning per-job sleeps and individual removals.

**Describe alternatives you've considered**

**Additional context**
Related: #1219 , #1314

Contributor guide

Open the contributing guide

Research direction

Start by tracing the existing clean_all_shuffle_data, clean_up_successful_job, and clean_up_failed_job flows, then review related issues #1219 and #1314. Identify where cleanup notifications and state.remove_job calls are dispatched before evaluating the shared remove_job_dir and batched remove_jobs boundaries. Done means cleanup is deduplicated, targeted to relevant executors, and dispatched in batches with tests for the resulting behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
backend, distributed-systems
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.