apache / apache/iceberg-python

Possible memory leak with to_arrow_batch_reader()

Open
#2,407 4 comments 2 reactions 0 assignees View on GitHub
stale
Dominant language
Python
Stars
1.1k
Forks
581
Avg merge
1d 17h
Merged PRs (30d)
78

Description

### Apache Iceberg version

0.9.1 (latest release)

### Please describe the bug 🐞

## Summary
It seems that there is memory leak when running to_arrow_batch_reader(), it takes ~30GB memory to read an iceberg table with single 40MB parquet files

Example code:
```python
import boto3
from pyiceberg.table import StaticTable, Table

def iceberg_table_from_metadata_path(metadata_path: str) -> StaticTable:
session = boto3.Session(region_name="")
credentials = session.get_credentials()
credentials = credentials.get_frozen_credentials()
table = StaticTable.from_metadata(
metadata_path,
{
"client.secret-access-key": credentials.secret_key,
"client.access-key-id": credentials.access_key,
"client.session-token": credentials.token,
"client.region": "",
},
)
return table

def main():
metadata_path = ".metadata.json"
iceberg_table = iceberg_table_from_metadata_path(metadata_path)
scan_kwargs = {"row_filter": f"PARTITION='train'"}
batch_reader = iceberg_table.scan(**scan_kwargs).to_arrow_batch_reader()
for batch in batch_reader:
print(f"Inside batch reader")

plan_files = iceberg_table.scan(**scan_kwargs).plan_files()
for file in plan_files:
print(file.file.file_path)

print("Hello from pyiceberg-test!")

if __name__ == "__main__":
main()

```

Running using memray
```
uv run memray run main.py
```

Charts

Image

Image

Dependencies
```
"boto3>=1.40.21",
"memray>=1.18.0",
"pyarrow>=21.0.0",
"pyiceberg==0.10.0rc1",
"s3fs>=0.4.2"
```

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the example in main.py with memray run main.py, focusing on the to_arrow_batch_reader() entry point and comparing memory during batch iteration with the later plan_files() call. Trace the scan and Arrow reader path to identify whether memory remains retained; done means the example reads the table without unbounded growth and the behavior is covered by an appropriate regression test.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.