apache / apache/iceberg-python

Possible memory leak with to_arrow_batch_reader()

Ouverte
#2,407 4 commentaires 2 réactions 0 personnes assignées Voir sur GitHub
stale
Langage dominant
Python
Étoiles
1.1k
Forks
581
Merge moyen
1 j 17 h
PR mergées (30 j)
77

Description

### Apache Iceberg version

0.9.1 (latest release)

### Please describe the bug 🐞

## Summary
It seems that there is memory leak when running to_arrow_batch_reader(), it takes ~30GB memory to read an iceberg table with single 40MB parquet files

Example code:
```python
import boto3
from pyiceberg.table import StaticTable, Table

def iceberg_table_from_metadata_path(metadata_path: str) -> StaticTable:
session = boto3.Session(region_name="")
credentials = session.get_credentials()
credentials = credentials.get_frozen_credentials()
table = StaticTable.from_metadata(
metadata_path,
{
"client.secret-access-key": credentials.secret_key,
"client.access-key-id": credentials.access_key,
"client.session-token": credentials.token,
"client.region": "",
},
)
return table

def main():
metadata_path = ".metadata.json"
iceberg_table = iceberg_table_from_metadata_path(metadata_path)
scan_kwargs = {"row_filter": f"PARTITION='train'"}
batch_reader = iceberg_table.scan(**scan_kwargs).to_arrow_batch_reader()
for batch in batch_reader:
print(f"Inside batch reader")

plan_files = iceberg_table.scan(**scan_kwargs).plan_files()
for file in plan_files:
print(file.file.file_path)

print("Hello from pyiceberg-test!")

if __name__ == "__main__":
main()

```

Running using memray
```
uv run memray run main.py
```

Charts

Image

Image

Dependencies
```
"boto3>=1.40.21",
"memray>=1.18.0",
"pyarrow>=21.0.0",
"pyiceberg==0.10.0rc1",
"s3fs>=0.4.2"
```

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Commencez par reproduire l’exemple de main.py avec memray run main.py, en vous concentrant sur le point d’entrée to_arrow_batch_reader() et en comparant la mémoire pendant l’itération par lots avec l’appel ultérieur à plan_files(). Suivez le chemin du scan et du lecteur Arrow pour déterminer si la mémoire reste retenue ; le travail est considéré comme terminé lorsque l’exemple lit la table sans croissance illimitée et que le comportement est couvert par un test de régression approprié.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
data-engineering, databases
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Active
Clarté
Plutôt claire
Accessibilité débutants
52/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.