apache / apache/hudi

[to be discussed] Users should be able to automatically block clustering if there are too many uncompacted/unarchived instants

Open
#17,903 2 comments 0 reactions 0 assignees View on GitHub
type:feature
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

### Feature Description

**What the feature achieves:**
Add clustering configs that prevent `scheduleClustering` from creating a new plan if
- There are too many unarchived/uncompacted instants in data table
- There are too many uncompacted instants in metadata table (MDT)

**Why this feature is needed:**
- If too many writes on a dataset are accumulated before clean and archival are executed on a dataset again, then the internal timeline may have thousands of instants. Additionally, clustering will cause the dataset partitions to contain replaced uncleaned older file groups, potentially leading to an uncleaned files building up.
- HUDI datasets need to undergo compaction on MDT once there have been enough writes accumulated. Any write or table service operation on the data table will cause a write on the metadata table. Delays in compaction will cause all writers to take more time in building their internal filesystem view when reading the metadata table, due to having to processing all uncompacted files. There is also an indirect impact of lock contention: clustering and write operations need to create a filesystem view while the lock is acquired here (in `org.apache.hudi.client.BaseHoodieTableServiceClient#scheduleTableServiceInternal`) - this will increase the time the lock is held and can cause other concurrent writers to get delayed/fail due to waiting for the table lock.

### User Experience

**How users will use this feature:**
- Configuration changes needed
- API changes
- Usage examples

### Hudi RFC Requirements

**RFC PR link:** (if applicable)

**Why RFC is/isn't needed:**
- Does this change public interfaces/APIs? (Yes/No)
- Does this change storage format? (Yes/No)
- Justification:

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading scheduleClustering and org.apache.hudi.client.BaseHoodieTableServiceClient#scheduleTableServiceInternal, focusing on how clustering plans and filesystem views are scheduled. The issue needs an agreed configuration and API design for data-table and metadata-table thresholds, usage examples, and corresponding validation before implementation can be considered done.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.