[MPP/Optimizer] make Optimizer smart enough to process complex quiries like multi-shuffle join.
Open
Nobody has claimed this yet.
priority/P1
- Dominant language
- C++
- Stars
- 1k
- Forks
- 423
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 24
Description
- recognize the multi joins with equivalent join keys. like a = b = c = d then the group by key is a
- choose shuffle key smartly. Shuffle keys can be the subset of join keys, but might result in data skew.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the MPP/Optimizer handling for multi-join plans and review how equivalent join keys and shuffle keys are represented. Use the a=b=c=d multi-join case as an initial example, then investigate how subset shuffle keys affect data skew. Done means the optimizer can process multi-shuffle joins and choose shuffle keys according to the stated constraints.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- databases, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100