opensearch-project / opensearch-project/data-prepper
[FEATURE] Group initial load tasks by file size in iceberg-source
@lawofcycles is already working on this.
Since Apr 7, 2026.
- Dominant language
- Java
- Stars
- 374
- Forks
- 354
- Avg merge
- 3d 18h
- Merged PRs (30d)
- 8
Description
Is your feature request related to a problem? Please describe.
The iceberg-source creates one task per data file during initial load. For tables with many small files, the coordination overhead per task (DynamoDB acquire/complete operations) can dominate the actual file processing time.
Describe the solution you'd like
Group multiple data files into a single initial load task based on total file size, consistent with the approach planned for SHUFFLE_WRITE tasks.
Additional context
Related: #6682 (source-layer shuffle), https://github.com/opensearch-project/data-prepper/issues/6724
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.