apache / apache/iceberg-python

Add maintenance action to remove dangling delete files

Abierto
#3,925 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
1.1k
Forks
581
Merge medio
1 d 17 h
PR fusionados (30 d)
77

Descripción

### Feature Request / Improvement

Streaming upsert writers (e.g. AWS Firehose Iceberg delivery) write equality-delete files on every commit. Compaction applies them into rewritten data files but leaves the entries in the manifests, and `expire_snapshots` can't touch files the current snapshot still references — so they accumulate without bound. On one of our production tables we measured ~90K dangling delete entries growing ~4.6K/day, and every query planning over recent partitions has to read the ever-growing delete manifests.

Java Iceberg handles this (`rewrite_data_files` with `remove-dangling-deletes`, `RemoveDanglingDeletesSparkAction`), but PyIceberg's `MaintenanceTable` currently only has `expire_snapshots`, and engines like Athena expose no statement for it either — so users on Athena/Firehose stacks have no non-Spark way out.

**Proposal:** `table.maintenance.remove_dangling_deletes()` — a metadata-only commit that:

- classifies per `(partition_spec_id, partition)`: an equality delete at sequence *s* is dangling iff no live data file in that partition has sequence < *s* (position deletes: <= *s*); ambiguous cases (unpartitioned specs, unknown content) are kept
- carries data manifests through unchanged, drops fully-dangling delete manifests, rewrites mixed ones to their surviving entries, and commits as a `replace` snapshot against the current ref

One enabler is worth a small standalone fix first: `ManifestWriterV2` hardcodes `content=data`, so PyIceberg currently can't write delete-content manifests at all.

We have a working implementation built on PyIceberg 0.12 internals (`write_manifest_list`, a `ManifestWriterV2` subclass, `AddSnapshotUpdate`/`SetSnapshotRefUpdate` with `AssertRefSnapshotId`), validated against production Glue/Athena tables, with a test matrix for the classification rules. Happy to contribute it if there's interest.

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Comienza con MaintenanceTable y ManifestWriterV2; después, inspecciona los elementos internos referenciados de write_manifest_list y de la actualización del snapshot. Usa las reglas de secuencia por partición indicadas y la matriz de pruebas de clasificación como criterios de aceptación; se considera terminado cuando un commit de reemplazo que solo modifica metadatos elimina únicamente los borrados colgantes y conserva las entradas ambiguas y activas.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
aws, python
Área
data-engineering, databases
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Activo
Claridad
Bastante claro
Aptitud para principiantes
45/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.