alibaba / alibaba/DataX

spark/flink hudi source with data reader or writer

Open
#2,320 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
17.4k
Forks
5.7k
PR merge metrics
No merged PRs in 30d

Description

datax 不能直接读写hudi数据,在企业系统中通过湖上建仓,这时候针对批量数据入湖,出湖缺少工具。
希望给datax 提供一个 hudi等湖 source/writer 能力。
但是datax 集成工作较大,希望通过spark中增加DataxRDD 来对datax reader/writer进行包装。实现打通各种湖组件。
也能够实现datax的分布式运行能力

Contributor guide

No contributing guide indexed for this repository

Research direction

No files or tests are named. Start by reviewing the existing DataX reader/writer architecture and the proposed Spark DataxRDD boundary, then define the scope for Hudi and other lake components. Done should mean DataX can run distributed batch reads and writes through the requested Spark/Flink integration.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.