apache / apache/datafusion-comet

Fork DataFusion SMJ in Comet repo

Open
#3,771 6 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Scala
Stars
1.3k
Forks
373
Avg merge
2d 4h
Merged PRs (30d)
198

Description

### What is the problem the feature request solves?

I would like to suggest that we fork the SMJ implementation from DataFusion and include it in the Comet repo.

I think this would let us customize/optimize for Spark use case and move faster without having to wait for DF release cycles. It also protects us from performance regressions when we don't have the time to test DF major releases until after they are released, which is often the case because we have to wait for iceberg-rust to upgrade first.

### Describe the potential solution

_No response_

### Additional context

_No response_

Contributor guide

Open the contributing guide

Research direction

Start by locating DataFusion's sort-merge join implementation and the corresponding Comet integration points. Review how iceberg-rust and DataFusion release compatibility affects the work. Done means the implementation is included in Comet and can be customized for Spark use cases without depending on future DataFusion releases.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust, scala
Domain
databases, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.