apache / apache/datafusion-ballista

Implement fast path for QueryStageExec when writing 1 shuffle partition

Open
#21 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Rust
Stars
2.1k
Forks
320
Avg merge
1d 22h
Merged PRs (30d)
66

Description

**Is your feature request related to a problem or challenge? Please describe what you are trying to do.**
As mentioned in https://github.com/apache/arrow-datafusion/pull/543#discussion_r650883972 we could optimize for the case where there is 1 output partition.

**Describe the solution you'd like**
Avoid the cost of computing hashes when output partition count is 1.

**Describe alternatives you've considered**
None

**Additional context**
None

Contributor guide

Open the contributing guide

Research direction

Start by locating QueryStageExec and tracing how it writes shuffle output when there is one partition. Read the optimization discussion in the linked DataFusion pull request for context, then run the relevant existing tests or benchmarks you find. Done means the one-partition path avoids hash computation while preserving correct shuffle output.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
data-engineering, distributed-systems
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.