opensearch-project / opensearch-project/data-prepper

[FEATURE] Group initial load tasks by file size in iceberg-source

Open
#6,725 0 comments 0 reactions 1 assignee View on GitHub

@lawofcycles is already working on this.

Since Apr 7, 2026.

enhancement
Dominant language
Java
Stars
374
Forks
354
Avg merge
3d 18h
Merged PRs (30d)
8

Description

Is your feature request related to a problem? Please describe.

The iceberg-source creates one task per data file during initial load. For tables with many small files, the coordination overhead per task (DynamoDB acquire/complete operations) can dominate the actual file processing time.

Describe the solution you'd like

Group multiple data files into a single initial load task based on total file size, consistent with the approach planned for SHUFFLE_WRITE tasks.

Additional context

Related: #6682 (source-layer shuffle), https://github.com/opensearch-project/data-prepper/issues/6724

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.