apache / apache/iceberg-python

Upsertion memory usage grows exponentially as table size grows

已關閉
#2,138 8 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
stale
主要語言
Python
星號
1.1k
分支
581
平均合併
1 天 13 小時
30 天內合併 PR
76

描述

### Apache Iceberg version

0.9.0

### Please describe the bug 🐞

Hello,

I am trying out the new upsert method on a table created as follows on the AWS Glue Catalog :

```python
catalog = load_catalog("glue", **{"type": "glue"})
schema = Schema(
NestedField(1, "dt_insert", StringType(), required=True),
NestedField(2, "controller_id", StringType(), required=True),
NestedField(3, "timestamp", StringType(), required=True),
NestedField(4, "data_type", StringType(), required=True),
NestedField(5, "parameter", StringType(), required=True),
NestedField(6, "value", StringType(), required=True),
NestedField(7, "data_source", StringType(), required=True),
NestedField(8, "unique_id", StringType(), required=True),
NestedField(9, "dt_import_utc", StringType(), required=True),
identifier_field_ids=[8], # 'unique_id' is the primary key
)
catalog.create_table(
identifier=.,
schema=schema,
partition_spec=PartitionSpec(
PartitionField(
source_id=9,
field_id=1000,
transform=IdentityTransform(),
name="dt_import_utc",
),
),
sort_order=SortOrder(
SortField(source_id=2), # controller_id
SortField(source_id=4), # data_type
SortField(source_id=3), # dt_insert
SortField(source_id=1), # timestamp
),
location="s3://bucket/prefix/catalog",
properties={
"write.format.default" : "parquet",
"write.target-file-size-bytes" : 134217728, # 128 MB
"write.metadata.delete-after-commit.enabled" : True,
"write.metadata.previous-versions-max" : 5
})
```

`unique_id` is a the merged string of columns 1, 2, 3, 4 and 5 with an underscore.

My upsertion is a simple pandas dataframe dumped into a dict and being transformed to pyarrow for upsertion (I know pyarrow accepts directly a pandas df but doing this for an internal reason that requires writing the df in a stage and read in another, so i am dumping as json for readability).

```python
iceberg_table = catalog.load_table(table_name)
table = pa.Table.from_pydict(to_upsert, schema=iceberg_table.schema().as_arrow())
iceberg_table.upsert(df=table, join_cols=["unique_id"])
```

With AWS Athena, a MERGE INTO statement takes about 3 seconds to run on the table and scans 1.11MB of data before completion.

```sql
MERGE INTO . target
USING . source
ON (target."unique_id" = source."unique_id")
WHEN MATCHED THEN
UPDATE SET "dt_insert" = source."dt_insert", "controller_id" = source."controller_id", "timestamp" = source."timestamp", "data_type" = source."data_type", "parameter" = source."parameter", "value" = source."value", "data_source" = source."data_source", "dt_import_utc" = source."dt_import_utc", "unique_id" = source."unique_id"
WHEN NOT MATCHED THEN
INSERT ("dt_insert", "controller_id", "timestamp", "data_type", "parameter", "value", "data_source", "dt_import_utc", "unique_id")
VALUES (source."dt_insert", source."controller_id", source."timestamp", source."data_type", source."parameter", source."value", source."data_source", source."dt_import_utc", source."unique_id")
```
(statement generated by using the `to_iceberg` from `awswrangler`)

Meanwhile when i try with pyiceberg upsert, it is using more than 10240MB. I am running on AWS lambda and it is causing an out of memory error.

I have no issue with the `append()` function, it completes fairly quickly but it seems that the upsert needs further optimization to be able to efficiently retrieve only relevant data.

Current table size is at 18.5GB for both Athena based statement and pyiceberg upsert function call.

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [ ] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [x] I cannot contribute a fix for this bug at this time

貢獻指南

這個儲存庫沒有索引到貢獻指南

研究方向

從 PyIceberg table.upsert 入口開始,使用提供的 AWS Glue 資料表 schema 和 join_cols=["unique_id"] 的重現來比較其與 append() 的行為。測量資料表持續成長時的記憶體使用量,並確認 upsert 是否能將資料擷取限制為相關資料列;當重現能在不出現所回報的記憶體不足錯誤的情況下完成時,即表示完成。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
aws, python
領域
databases
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
活躍
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。