Add source data estimators to optimize ingestion runs
Open
area:ingest
from-jira
priority:high
type:feature
- Dominant language
- Java
- Stars
- 6.2k
- Forks
- 2.5k
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 111
Description
Estimate how much new data is present to be ingested from a given data source, and schedule DeltaStreamer jobs based on that.
## JIRA info
- Link: https://issues.apache.org/jira/browse/HUDI-5650
- Type: New Feature
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the linked HUDI-5650 JIRA issue and the DeltaStreamer ingestion and scheduling entry points. Trace how a source is currently assessed before a job runs, then define the estimator and scheduling behavior from the issue requirements. Done means source data volume can guide DeltaStreamer job scheduling, with tests covering the new behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering, stream-processing
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100