apache / apache/datafusion-ballista
Implement fast path for QueryStageExec when writing 1 shuffle partition
- Dominant language
- Rust
- Stars
- 2.1k
- Forks
- 320
- Avg merge
- 1d 22h
- Merged PRs (30d)
- 66
Description
**Is your feature request related to a problem or challenge? Please describe what you are trying to do.**
As mentioned in https://github.com/apache/arrow-datafusion/pull/543#discussion_r650883972 we could optimize for the case where there is 1 output partition.
**Describe the solution you'd like**
Avoid the cost of computing hashes when output partition count is 1.
**Describe alternatives you've considered**
None
**Additional context**
None
Contributor guide
Research direction
Start by locating QueryStageExec and tracing how it writes shuffle output when there is one partition. Read the optimization discussion in the linked DataFusion pull request for context, then run the relevant existing tests or benchmarks you find. Done means the one-partition path avoids hash computation while preserving correct shuffle output.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- data-engineering, distributed-systems
- Issue type
- Feature
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100