apache / apache/hudi

there is no data when a couple of hudi tables join

Open
#10,366 4 comments 0 reactions 0 assignees View on GitHub
area:incr-processing priority:high
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

There is one etl job run every hour and it is insert overwrite one table from the results that is generated by some hudi table join. It happens like one a week that there is no data inserted.

Environment Description

Hudi version : 0.9.1

Spark version : 3.0.1

Hive version : 3

Hadoop version : 3.2.2

Storage (HDFS/S3/GCS..) : s3

Running on Docker? (yes/no) : no

what cab be the reason ? Is there any way to debug this kind of issues or how to get the more metrics for it?

Contributor guide

No contributing guide indexed for this repository

Research direction

No source files or tests are identified. Start by reproducing the hourly insert-overwrite job with the Hudi, Spark, Hive, Hadoop, and S3 versions listed, then inspect available job logs and metrics around the joined tables. Done means identifying a likely cause or documenting the evidence needed to diagnose the missing output.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.