apache / apache/iceberg-python

Upsert gets slow on tables with many columns

Open
#3,860 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.1k
Forks
581
Avg merge
1d 17h
Merged PRs (30d)
78

Description

### Feature Request / Improvement

`upsert` compares the matched rows one cell at a time, so it gets slower with every extra column, not just with every extra row.

On a table with 200 columns, comparing 20k matched rows takes around 20 seconds on my machine, before anything is written.

To reproduce:

```python
import time
import pyarrow as pa
from pyiceberg.table.upsert_util import get_rows_to_update

rows, cols = 20_000, 200
table = pa.table({"pk": pa.array(range(rows)), **{f"c{i}": pa.array([float(i)] * rows) for i in range(cols)}})

start = time.monotonic()
get_rows_to_update(table, table, ["pk"]) # nothing has changed
print(time.monotonic() - start)
```

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with pyiceberg.table.upsert_util.get_rows_to_update and run the provided PyArrow reproduction with 20,000 rows and 200 columns. Compare the current behavior on unchanged tables and trace where matched rows are compared; done means the no-change upsert comparison avoids the reported column-dependent slowdown while preserving the expected rows to update.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
databases, performance
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.