apache / apache/iceberg-python

Metadata inspection APIs fail with struct.error after int→long / float→double type promotion

オープン
#3,744 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
1.1k
フォーク
581
平均マージ
1日 17時間
マージ済み PR(30日)
77

説明

### Apache Iceberg version

0.11.0 (latest release)

### Please describe the bug 🐞

### Summary

After a type promotion that the spec allows (`int` → `long`, `float` → `double`), **all metadata inspection APIs raise `struct.error`** on any table that already contains data files written before the promotion:

| Promotion | `inspect.files()` | `inspect.entries()` | `inspect.data_files()` | `inspect.all_files()` |
|---|---|---|---|---|
| `int` → `long` | ❌ | ❌ | ❌ | ❌ |
| `float` → `double` | ❌ | ❌ | ❌ | ❌ |

```
struct.error: unpack requires a buffer of 8 bytes
```

Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.

### Reproduction

Self-contained, no cloud services or network required:

```python
import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType

warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")

tbl = catalog.create_table("ns.t", schema=Schema(
NestedField(1, "name", StringType(), required=False),
NestedField(2, "qty", IntegerType(), required=False),
))

# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
[{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
schema=pa.schema([pa.field("name", pa.string(), nullable=True),
pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")

# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
update.update_column("qty", field_type=LongType())

tbl = catalog.load_table("ns.t")
try:
tbl.inspect.files()
except Exception:
traceback.print_exc()

shutil.rmtree(warehouse, ignore_errors=True)
```

**Output**

```
before promotion: 1 row(s) -> OK
Traceback (most recent call last):
...
File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
"lower_bound": from_bytes(field.field_type, lower_bound)
File ".../pyiceberg/conversions.py", line 337, in _
return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes
```

Replacing `IntegerType()/LongType()/pa.int32()` with `FloatType()/DoubleType()/pa.float32()` reproduces the same failure, as does calling `entries()`, `data_files()` or `all_files()` instead of `files()`.

### Expected behavior

`inspect.files()` and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.

### Analysis

`InspectTable._get_files_from_manifest` decodes `lower_bounds` / `upper_bounds` via `from_bytes(field.field_type, ...)`, i.e. using the **current** schema type.

Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after `int` → `long`, pre-existing files still carry 4-byte bounds while the current field type is `long`, and `_LONG_STRUCT.unpack` (8 bytes) fails.

As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the **encoded byte length** rather than trusting the current schema type.

### Workaround

`Table.scan().plan_files()` returns `DataFile` objects without decoding bounds, so file-level metadata (`file_path`, `record_count`, `sort_order_id`, `spec_id`, …) remains reachable:

```python
tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
```

### Environment

- pyiceberg 0.11.1 (latest release)
- pyarrow 25.0.0, SQLAlchemy 2.0.51
- Python 3.14.5, macOS (arm64)

Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose `int` column had been promoted to `bigint` via Athena; the reproduction above shows it is catalog-independent.

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

自己完結型スクリプトで失敗を再現し、その後 pyiceberg/table/inspect.py の InspectTable._get_files_from_manifest と pyiceberg/conversions.py の from_bytes を読みます。int→long と float→double について、エンコードされた bound の長さを現在のフィールド型と比較します。files()、entries()、data_files()、all_files() が struct.error を発生させずに bounds または None を返せば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
databases
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
静か
明瞭さ
明確に書かれている
初心者へのやさしさ
68/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。