apache / apache/iceberg-python

load_table consumes enormous amounts of memory on large metadata file

Abierto
#3,162 2 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
1.1k
Forks
581
Merge medio
1 d 17 h
PR fusionados (30 d)
78

Descripción

### Apache Iceberg version

0.11.0 (latest release)

### Please describe the bug 🐞

Apologies, this is a bit of a fuzzy one right now, but I thought reporting it anyway.

**Context:**
We're using Iceberg with AWS Glue and AWS S3 as storage. In S3 there are roughly speaking 3 kinds of files (metadata, manifests, and data files). The first one that is read when loading a table via `catalog.load_table()` is the metadata file. The metadata file contains information on all current* snapshots and schema versions of the table. py-iceberg seems to load these completely into memory.

**Issue:**
As we worked on the Iceberg table, there were a lot of snapshots created over time and with that a lot of schema versions. This led to the latest metadata file to be grow to ~10MB gzip compressed (or ~250MB uncompressed JSON). When we load this table via `catalog.load_table()` it consumes ~4GB of memory (total usage of the python process in memray). This is a lot - especially since we only need the latest snapshot and the respective schema version. (Which is probably true for most users I guess.)

**Semi-Workaround:**
One could try to expire some snapshots, e.g. via Sparks `expire_snapshots` procedure [https://iceberg.apache.org/docs/1.10.0/spark-procedures/#expire_snapshots], but it will not get rid of the old / unused schemas unless you set `clean_expired_metadata` as well (which is only supported since 1.10.x, so relatively new).

**(Preliminary) Root-Cause:**
I believe the issue is that we leverage Pydantic's `model_validate_json` in https://github.com/apache/iceberg-python/blob/44ce51a939ccbacf9c87ce6593ad43a752b0871b/pyiceberg/table/metadata.py#L663, which loads the whole JSON into memory and then we seem to keep the full `TableMetadata` object around.

**Suggestion:**
Would it make sense to parse the JSON not fully into memory and load the needed snapshots and schemas lazy / on demand? (Would be also fine, if that is a configurable option of `catalog.load_table()`)

**Remark:**
Obviously we could blame this on an un-maintained Iceberg table, but I think it would be good for the pyIceberg lib to be robust against such scenarios, hence why I opened the issue.

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [ ] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Comienza en pyiceberg/table/metadata.py alrededor de la línea 663, donde catalog.load_table() utiliza model_validate_json de Pydantic, y reproduce el uso máximo de memoria con un archivo de metadatos grande usando memray. Investiga cómo se conservan los snapshots y los schemas durante la carga; la tarea estará terminada cuando los archivos de metadatos grandes se carguen con un uso máximo de memoria sustancialmente menor, mientras el snapshot y el schema más recientes sigan estando disponibles.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
aws, python
Área
data-engineering, databases
Tipo de issue
Error
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Tranquilo
Claridad
Necesita aclaración
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.