apache / apache/iceberg

RemoveDanglingDeleteFiles does not support branches and skips unpartitioned tables

Open
#15,369 1 comment 1 reaction 0 assignees View on GitHub
improvement
Dominant language
Java
Stars
9.2k
Forks
3.5k
Avg merge
2d 16h
Merged PRs (30d)
129

Description

### Feature Request / Improvement

The current implementation of rewrite data files with `removeDanglingDeletes` only works on the main branch of the iceberg table.

The RemoveDanglingDeleteFiles action has two issues:

1. Unpartitioned tables are silently skipped

The execute() method returns early with an empty result for unpartitioned tables:

```
if (table.specs().size() == 1 && table.spec().isUnpartitioned()) {
return ImmutableRemoveDanglingDeleteFiles.Result.builder()
.removedDeleteFiles(Collections.emptyList())
.build();
}
```

2. No branch support

The action always operates on the main branch. There is no API to target a specific branch, and the Spark implementation reads metadata tables without branch scoping. This also means that when RewriteDataFilesSparkAction invokes RemoveDanglingDeleteFiles internally (via the remove-dangling-deletes option), it ignores the branch that the rewrite is targeting.

Proposed Changes

API (RemoveDanglingDeleteFiles):
- Add a toBranch(String branch) method to allow targeting a specific branch.

**Background:**
- We're leveraging Flink to write to iceberg in streaming fashion, but using Write-Audit-Publish pattern. So flink writes to a branch.
- Periodic Spark job that reads the latest changes on the branch, runs audit and tries to merge/fast-forward to main.
- Sync to Snowflake.

Since Flink streaming data in Upsert mode generates `equality-deletes` on branch, it is also present in the metadata on main.

Snowflake however doesn't support equality deletes for managed tables and this requires us to remove equality deletes from main branch, which is why we need the capability to remove dangling deletes from the branch, before we fast-forward to main.

### Query engine

Spark

### Willingness to contribute

- [ ] I can contribute this improvement/feature independently
- [x] I would be willing to contribute this improvement/feature with guidance from the Iceberg community
- [ ] I cannot contribute this improvement/feature at this time

Contributor guide

Open the contributing guide

Research direction

Start at RemoveDanglingDeleteFiles.execute() and its current unpartitioned-table early return, then trace how RewriteDataFilesSparkAction invokes it. Review the Spark metadata-table reads and branch-related action APIs. Done means a toBranch(String branch) API works for partitioned and unpartitioned tables and the rewrite action scopes dangling-delete processing to its target branch.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, databases
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.