[DISCUSSION] JOIN "task force" / project team
- Dominant language
- Rust
- Stars
- 9.3k
- Forks
- 2.4k
- Avg merge
- 3d 7h
- Merged PRs (30d)
- 344
Description
# What I see (what problem we are trying to solve)
DataFusion's current join implementations are fairly basic. They are functional enough to run TPCH and TPC-DS, but lack other features such as larger-than-memory processing, ASOF joins, complete subquery support and more.
There seems to be a non trivial desire in the community to improve this.
Some examples of issues / tickets related to enhanced join support / features:
## Subqueries (which are implemented as joins)
- [ ] https://github.com/apache/datafusion/issues/5483
- [ ] https://github.com/apache/datafusion/issues/5492
- [ ] https://github.com/apache/datafusion/issues/14554
## Join Features
- [ ] https://github.com/apache/datafusion/issues/12454
- [ ] https://github.com/apache/datafusion/issues/15784
- [ ] https://github.com/apache/datafusion/issues/14239
- [ ] https://github.com/apache/datafusion/issues/14238
- [ ] https://github.com/apache/datafusion/issues/13765
- [ ] https://github.com/apache/datafusion/issues/13003
- [x] https://github.com/apache/datafusion/issues/12952
- [x] https://github.com/apache/datafusion/issues/10048
- [ ] https://github.com/apache/datafusion/issues/3843
## Specialized Joins
- [x] https://github.com/apache/datafusion/issues/9846
- [ ] https://github.com/apache/datafusion/issues/318
- [x] https://github.com/apache/datafusion/issues/13471
- [ ] https://github.com/apache/datafusion/issues/13232
- [ ] https://github.com/apache/datafusion/issues/13181
- [x] https://github.com/apache/datafusion/issues/13138
## Performance
- [ ] https://github.com/apache/datafusion/issues/15382
- [x] https://github.com/apache/datafusion/issues/7955
- [ ] https://github.com/apache/datafusion/issues/14758
- [ ] https://github.com/apache/datafusion/issues/13620
# What is blocking significant forward progress
In my mind, the major challenge is that "improving" `JOIN`s can get arbitrarily complicated. There are dozens of academic paper each year on various aspects of join implemnetations, and designing / implementing join capabilities is a substantial engineering effort.
I spent 6 years of my life doing joins at Vertica where they accounted for around 50% of the optimizer's complexity, to give some sense
I don't think the issue is that any particular feature is super complicated to understand, but defining the overall goal, the framework that will accomodate the goal, and then breaking it down into implementable pieces itself I think will require both specialized knowledge and substantial time.
## What I suggest
I suggest that people with the relevant skills and time to invest gather together to drive this process worward
1. plan out a "join roadmap" (aka prioritize what join features they will push forward)
2. Figure out what, if any, new structures are in place
3. Start breaking it down into smaller tickets
I can't personally lead such an effort, but I am filing this ticket to try and help connect the relevant people in the community that can.
Some potential people that could help (sorry if I didn't list you)
* @duongcongtoai -- the discussion on https://github.com/apache/datafusion/issues/14554#issuecomment-2798943345
* @xudong963 who has experience in this area
* @Dandandan @comphead and @korowa who contributed substantially to the existing joins
* @mingmwang and @jackwener who contributed significantly to the original subquery implementation
* @liukun4515 who likewise helped significantly
* @suibianwanwank
## Related content:
Related blogs (join ordering section in part 2): https://www.influxdata.com/blog/optimizing-sql-dataframes-part-two/
Contributor guide
Research direction
Start by reviewing the linked subquery, join-feature, specialized-join, and performance issues, then read the related join-ordering blog section. The discussion calls for a prioritized join roadmap, assessment of existing structures, and smaller implementation tickets; it does not name specific files or tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust, sql
- Domain
- databases
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100