[UMBRELLA] Support Apache Calcite for writing/querying Hudi datasets
- Dominant language
- Java
- Stars
- 6.2k
- Forks
- 2.5k
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 111
Description
(More details to be added)
## JIRA info
- Link: https://issues.apache.org/jira/browse/HUDI-1387
- Type: Epic
---
## Comments
10/Nov/20 20:32;xushiyan;[~vinoth] Made this under presto integration component. Shall we rename it to "Query Engine Support"?;;;
---
12/Nov/20 03:42;vinoth;I will leave this in Common core for now. Probably worth having the presto specific component as-is. We can create a new Calcite specific one down the line. ;;;
---
12/Nov/20 23:35;xushiyan;[~vinoth] ok sounds good.;;;
---
21/Dec/21 17:37;vinoth; The main idea here was that we introduce our own SQL using Calcite.. Biggest question in my mind is. Could we just get by using Flink, given it already is built on Calcite. Is Flink SQL extensible ?;;;
---
21/Dec/21 18:27;xushiyan;[~x1q1j1] can you share some thoughts on this pls? ;;;
---
23/Dec/21 11:35;x1q1j1;This sounds very good, I think about it, there are probably two big things that need to be made, one is that the parser can support the commonly used olap or engine such as flink/spark/presto. One is to convert the parser into Hudi's own execution plan.;;;
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue provides no files, tests, or entry points; start by reviewing the linked JIRA epic and existing query-engine integrations. Clarify whether Hudi should use Flink SQL/Calcite or define its own SQL parser and execution plan, then document the agreed scope before implementation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering, databases
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100