apache / apache/iceberg-python
Null list<struct<...>> is written and read as an empty list
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 1.1k
- 分支
- 588
- 平均合併
- 1 天 23 小時
- 30 天內合併 PR
- 84
描述
### Apache Iceberg version
main (development)
### Please describe the bug 🐞
A null `list>` is silently rebuilt as an **empty** list. This is known on
the read path — `tests/integration/test_reads.py::test_null_list_and_map` currently
asserts the corrupted value, with the correct assertion commented out pending
apache/arrow#38809:
```python
# This should be:
# assert arrow_table["col_list_with_struct"].to_pylist() == [None, [{'test': 1}]]
# Once https://github.com/apache/arrow/issues/38809 has been fixed
assert arrow_table["col_list_with_struct"].to_pylist() == [[], [{"test": 1}]]
```
Two things seem worth reporting on top of that.
**It also affects the write path**, where the consequence is worse. The Parquet file
pyiceberg writes contains an empty list, so the null is gone at rest and no reader —
pyiceberg, Spark, Trino — can recover it. On read the file is at least still correct.
**It does not depend on the upstream Arrow fix.** `pa.LargeListArray.from_arrays`
takes a `mask` argument — since well before pyiceberg's `pyarrow>=18.0.0` floor — so
this particular null loss can be fixed independently of apache/arrow#38809, which is
still open.
This is the `array>` case from #251. That issue was closed in March 2025 on
the strength of this test existing, but the assertion it makes is the corrupted one;
the `array` case in the issue body was genuinely fixed by #252, while the
`array>` case in the issue *title* — which @HonahX flagged as
remaining broken in
https://github.com/apache/iceberg-python/pull/252#discussion_r1467065763 — was not.
#### Reproduction (write path)
pyiceberg 0.11.1, pyarrow 25.0.1:
```python
import os, shutil, glob
import pyarrow as pa, pyarrow.parquet as pq
from pyiceberg.catalog.sql import SqlCatalog
WH = "/tmp/wh"; shutil.rmtree(WH, ignore_errors=True); os.makedirs(WH)
sch = pa.schema([
pa.field("id", pa.int32(), nullable=False),
pa.field("l_struct", pa.list_(pa.field("element", pa.struct([pa.field("x", pa.int32())]), nullable=True)), nullable=True),
pa.field("l_int", pa.list_(pa.field("element", pa.int32(), nullable=True)), nullable=True),
])
tbl = pa.table({"id": [1, 2, 3, 4],
"l_struct": [[{"x": 1}], [], None, [{"x": 3}]],
"l_int": [[1], [], None, [3]]}, schema=sch)
cat = SqlCatalog("r", uri=f"sqlite:///{WH}/c.db", warehouse=f"file://{WH}")
cat.create_namespace("ns")
it = cat.create_table("ns.t", schema=tbl.schema)
it.append(tbl)
out = it.scan().to_arrow()
for c in ("l_struct", "l_int"):
print(f"{c:9s} in={tbl.column(c).to_pylist()!s:35s} out={out.column(c).to_pylist()}")
# the loss is already in the file on disk, not in the read path
f = glob.glob(f"{WH}/**/*.parquet", recursive=True)[0]
print("raw parquet:", pq.read_table(f).column("l_struct").to_pylist())
```
Output:
```
l_struct in=[[{'x': 1}], [], None, [{'x': 3}]] out=[[{'x': 1}], [], [], [{'x': 3}]]
l_int in=[[1], [], None, [3]] out=[[1], [], None, [3]]
raw parquet: [[{'x': 1}], [], [], [{'x': 3}]]
```
`l_int` round-trips correctly, and writing the same `pa.Table` with `pq.write_table`
preserves the null, so the loss is not pyarrow's.
#### Cause
`ArrowProjectionVisitor.list` rebuilds the array when the element is a struct
(`pyiceberg/io/pyarrow.py:2078` on `main` @ 7539661):
```python
if isinstance(value_array, pa.StructArray):
# This can be removed once this has been fixed:
# https://github.com/apache/arrow/issues/38809
list_array = pa.LargeListArray.from_arrays(list_array.offsets, value_array)
```
`from_arrays` receives the offsets buffer alone, which cannot express a null list, and
no mask, so the validity bitmap is dropped. That is also why only this one shape is
affected: the `struct` visitor passes `mask=struct_array.is_null()`, the `map` visitor
does not rebuild at all, and a list whose element is a primitive never enters this
branch. The visitor runs on both paths, which is why the same root cause shows up as
the read-side assertion above and as the write-side corruption here.
#### Fix
Carrying the mask over is enough:
```python
list_array = pa.LargeListArray.from_arrays(list_array.offsets, value_array, mask=list_array.is_null())
```
With that change the reproduction above returns `None` for both columns, and
`test_null_list_and_map` passes with its commented-out assertion restored. I have not
looked at int32-offset or sliced-array handling of this call, which the existing line
already relies on; that appears independent of the mask.
I have this on a branch with a unit test covering the write path and the integration
assertion un-commented, and can open a PR.
Found while testing a third-party Iceberg writer against pyiceberg as a reader.
### Willingness to contribute
- [x] I can contribute a fix for this bug independently
貢獻指南
這個儲存庫沒有索引到貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 pyiceberg/io/pyarrow.py:2078 的 ArrowProjectionVisitor.list 開始,檢查 tests/integration/test_reads.py::test_null_list_and_map。確認該重現能在寫入的 Parquet 檔案中以及讀取時保留 null struct lists,然後執行聚焦的整合測試和相關的單元測試;當恢復的 assertion 通過且未使 primitive-list 處理出現回歸時,即表示完成。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- data-engineering, databases
- Issue 類型
- 缺陷
- 難度
- 2/5
- 預估耗時
- 半天
- 活躍度
- 活躍
- 描述清晰度
- 描述清楚
- 新手友好度
- 45/100