apache / apache/hudi

Add source data estimators to optimize ingestion runs

Open
#15,722 0 comments 0 reactions 0 assignees View on GitHub
area:ingest from-jira priority:high type:feature
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

Estimate how much new data is present to be ingested from a given data source, and schedule DeltaStreamer jobs based on that.

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-5650
- Type: New Feature

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the linked HUDI-5650 JIRA issue and the DeltaStreamer ingestion and scheduling entry points. Trace how a source is currently assessed before a job runs, then define the estimator and scheduling behavior from the issue requirements. Done means source data volume can guide DeltaStreamer job scheduling, with tests covering the new behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering, stream-processing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.