apache / apache/iceberg-python

Add maintenance action to remove dangling delete files

Aberta
#3,925 0 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
1.1k
Forks
581
Merge médio
1d 13h
PRs com merge (30d)
76

Descrição

### Feature Request / Improvement

Streaming upsert writers (e.g. AWS Firehose Iceberg delivery) write equality-delete files on every commit. Compaction applies them into rewritten data files but leaves the entries in the manifests, and `expire_snapshots` can't touch files the current snapshot still references — so they accumulate without bound. On one of our production tables we measured ~90K dangling delete entries growing ~4.6K/day, and every query planning over recent partitions has to read the ever-growing delete manifests.

Java Iceberg handles this (`rewrite_data_files` with `remove-dangling-deletes`, `RemoveDanglingDeletesSparkAction`), but PyIceberg's `MaintenanceTable` currently only has `expire_snapshots`, and engines like Athena expose no statement for it either — so users on Athena/Firehose stacks have no non-Spark way out.

**Proposal:** `table.maintenance.remove_dangling_deletes()` — a metadata-only commit that:

- classifies per `(partition_spec_id, partition)`: an equality delete at sequence *s* is dangling iff no live data file in that partition has sequence < *s* (position deletes: <= *s*); ambiguous cases (unpartitioned specs, unknown content) are kept
- carries data manifests through unchanged, drops fully-dangling delete manifests, rewrites mixed ones to their surviving entries, and commits as a `replace` snapshot against the current ref

One enabler is worth a small standalone fix first: `ManifestWriterV2` hardcodes `content=data`, so PyIceberg currently can't write delete-content manifests at all.

We have a working implementation built on PyIceberg 0.12 internals (`write_manifest_list`, a `ManifestWriterV2` subclass, `AddSnapshotUpdate`/`SetSnapshotRefUpdate` with `AssertRefSnapshotId`), validated against production Glue/Athena tables, with a test matrix for the classification rules. Happy to contribute it if there's interest.

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Direção de pesquisa

Comece com MaintenanceTable e ManifestWriterV2 e, em seguida, inspecione os componentes internos referenciados de write_manifest_list e da atualização do snapshot. Use as regras de sequência por partição especificadas e a matriz de testes de classificação como critérios de aceitação; considera-se concluído quando um commit de substituição somente de metadados remove apenas as exclusões pendentes, preservando as entradas ambíguas e ativas.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
aws, python
Domínio
data-engineering, databases
Tipo de issue
Funcionalidade
Dificuldade
5/5
Tempo estimado
Mais de uma semana
Status de atividade
Ativa
Clareza
Razoavelmente clara
Facilidade para iniciantes
45/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.