apache / apache/iceberg-python

Possible memory leak with to_arrow_batch_reader()

Aberta
#2,407 4 comentários 2 reações 0 responsáveis Ver no GitHub
stale
Linguagem predominante
Python
Estrelas
1.1k
Forks
581
Merge médio
1d 13h
PRs com merge (30d)
76

Descrição

### Apache Iceberg version

0.9.1 (latest release)

### Please describe the bug 🐞

## Summary
It seems that there is memory leak when running to_arrow_batch_reader(), it takes ~30GB memory to read an iceberg table with single 40MB parquet files

Example code:
```python
import boto3
from pyiceberg.table import StaticTable, Table

def iceberg_table_from_metadata_path(metadata_path: str) -> StaticTable:
session = boto3.Session(region_name="")
credentials = session.get_credentials()
credentials = credentials.get_frozen_credentials()
table = StaticTable.from_metadata(
metadata_path,
{
"client.secret-access-key": credentials.secret_key,
"client.access-key-id": credentials.access_key,
"client.session-token": credentials.token,
"client.region": "",
},
)
return table

def main():
metadata_path = ".metadata.json"
iceberg_table = iceberg_table_from_metadata_path(metadata_path)
scan_kwargs = {"row_filter": f"PARTITION='train'"}
batch_reader = iceberg_table.scan(**scan_kwargs).to_arrow_batch_reader()
for batch in batch_reader:
print(f"Inside batch reader")

plan_files = iceberg_table.scan(**scan_kwargs).plan_files()
for file in plan_files:
print(file.file.file_path)

print("Hello from pyiceberg-test!")

if __name__ == "__main__":
main()

```

Running using memray
```
uv run memray run main.py
```

Charts

Image

Image

Dependencies
```
"boto3>=1.40.21",
"memray>=1.18.0",
"pyarrow>=21.0.0",
"pyiceberg==0.10.0rc1",
"s3fs>=0.4.2"
```

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Direção de pesquisa

Comece reproduzindo o exemplo em main.py com memray run main.py, concentrando-se no ponto de entrada to_arrow_batch_reader() e comparando a memória durante a iteração dos lotes com a chamada posterior a plan_files(). Rastreie o caminho de scan e do leitor Arrow para identificar se a memória continua retida; considera-se concluído quando o exemplo lê a tabela sem crescimento ilimitado e o comportamento está coberto por um teste de regressão apropriado.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
python
Domínio
data-engineering, databases
Tipo de issue
Bug
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Ativa
Clareza
Razoavelmente clara
Facilidade para iniciantes
52/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.