apache / apache/iceberg-python

Upsert fails after adding in a new column, target field not found

未关闭
#2,467 6 条评论 6 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
1.1k
派生
581
平均合并
1 天 17 小时
30 天内合并 PR
77

描述

### Apache Iceberg version

0.10.0 (latest release)

### Please describe the bug 🐞

The upsert works perfectly fine until I needed to add a new field in the table.

Add the new column in the table
```python
# Add in a column to an existing table

from pyiceberg.types import TimestamptzType, TimestampType

table = catalog.load_table(table_identifier)

(
table.update_schema()
.add_column("created_at", TimestamptzType(), doc="UTC created time", required=False)
.commit()
)

print("New schema:", table.schema())
```

Upsert the records
```python
# Batch the records in 1000s
for rb in arrow_table_fixed.to_batches(max_chunksize=1000):
batch_tbl = pa.Table.from_batches([rb])

# Upsert the data into the Iceberg table
try:
upd = iceberg_table.upsert(batch_tbl)
print("Upserted data into the Iceberg table.")
print(upd)
except Exception as e:
print(f"An error occurred during upsert: {e}")
```

Error message saying that the target schema doesn't have the new column
```error
An error occurred during upsert: Target schema's field names are not matching the table's field names: ['cik_str', 'ticker', 'title', 'created_at'], ['cik_str', 'ticker', 'title']
```

Checked the target schema on Iceberg and the column is definitely there
```python
# Get the schema from the Iceberg table
iceberg_table = catalog.load_table(table_identifier)
# 2) Get the PyArrow schema directly from the Iceberg schema
arrow_schema = iceberg_table.schema().as_arrow()
print(arrow_schema.schema)
```

output
```
cik_str: large_string not null
-- field metadata --
PARQUET:field_id: '1'
ticker: large_string not null
-- field metadata --
PARQUET:field_id: '2'
title: large_string
-- field metadata --
PARQUET:field_id: '3'
created_at: timestamp[us, tz=UTC]
-- field metadata --
doc: 'UTC created time'
PARQUET:field_id: '5'
```

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [ ] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [x] I cannot contribute a fix for this bug at this time

贡献指南

这个仓库没有索引到贡献指南

调研方向

首先,使用 table.update_schema()、catalog.load_table() 和 iceberg_table.upsert(),并添加 created_at 字段,重现所提供的序列。跟踪 upsert 目标 schema 验证,该验证会报告字段名称不匹配。完成标准是:添加列后,upsert 成功且不再出现目标 schema 错误,同时保留所演示的 schema 行为。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
data-engineering, databases
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
活跃
描述清晰度
基本清楚
新手友好度
55/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。