apache / apache/datafusion-comet
Fork DataFusion SMJ in Comet repo
- Dominant language
- Scala
- Stars
- 1.3k
- Forks
- 373
- Avg merge
- 2d 4h
- Merged PRs (30d)
- 198
Description
### What is the problem the feature request solves?
I would like to suggest that we fork the SMJ implementation from DataFusion and include it in the Comet repo.
I think this would let us customize/optimize for Spark use case and move faster without having to wait for DF release cycles. It also protects us from performance regressions when we don't have the time to test DF major releases until after they are released, which is often the case because we have to wait for iceberg-rust to upgrade first.
### Describe the potential solution
_No response_
### Additional context
_No response_
Contributor guide
Research direction
Start by locating DataFusion's sort-merge join implementation and the corresponding Comet integration points. Review how iceberg-rust and DataFusion release compatibility affects the work. Done means the implementation is included in Comet and can be customized for Spark use cases without depending on future DataFusion releases.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust, scala
- Domain
- databases, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 32/100