spark/flink hudi source with data reader or writer
Open
- Dominant language
- Java
- Stars
- 17.4k
- Forks
- 5.7k
- PR merge metrics
- No merged PRs in 30d
Description
datax 不能直接读写hudi数据,在企业系统中通过湖上建仓,这时候针对批量数据入湖,出湖缺少工具。
希望给datax 提供一个 hudi等湖 source/writer 能力。
但是datax 集成工作较大,希望通过spark中增加DataxRDD 来对datax reader/writer进行包装。实现打通各种湖组件。
也能够实现datax的分布式运行能力
Contributor guide
No contributing guide indexed for this repository
Research direction
No files or tests are named. Start by reviewing the existing DataX reader/writer architecture and the proposed Spark DataxRDD boundary, then define the scope for Hudi and other lake components. Done should mean DataX can run distributed batch reads and writes through the requested Spark/Flink integration.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, spark
- Domain
- data-engineering, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100