apache / apache/iceberg-python

Metadata inspection APIs fail with struct.error after int→long / float→double type promotion

Đang mở
#3,744 3 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
1.1k
Fork
581
Merge trung bình
1 ngày 17 giờ
Pull request đã merge (30 ngày)
78

Mô tả

### Apache Iceberg version

0.11.0 (latest release)

### Please describe the bug 🐞

### Summary

After a type promotion that the spec allows (`int` → `long`, `float` → `double`), **all metadata inspection APIs raise `struct.error`** on any table that already contains data files written before the promotion:

| Promotion | `inspect.files()` | `inspect.entries()` | `inspect.data_files()` | `inspect.all_files()` |
|---|---|---|---|---|
| `int` → `long` | ❌ | ❌ | ❌ | ❌ |
| `float` → `double` | ❌ | ❌ | ❌ | ❌ |

```
struct.error: unpack requires a buffer of 8 bytes
```

Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.

### Reproduction

Self-contained, no cloud services or network required:

```python
import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType

warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")

tbl = catalog.create_table("ns.t", schema=Schema(
NestedField(1, "name", StringType(), required=False),
NestedField(2, "qty", IntegerType(), required=False),
))

# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
[{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
schema=pa.schema([pa.field("name", pa.string(), nullable=True),
pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")

# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
update.update_column("qty", field_type=LongType())

tbl = catalog.load_table("ns.t")
try:
tbl.inspect.files()
except Exception:
traceback.print_exc()

shutil.rmtree(warehouse, ignore_errors=True)
```

**Output**

```
before promotion: 1 row(s) -> OK
Traceback (most recent call last):
...
File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
"lower_bound": from_bytes(field.field_type, lower_bound)
File ".../pyiceberg/conversions.py", line 337, in _
return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes
```

Replacing `IntegerType()/LongType()/pa.int32()` with `FloatType()/DoubleType()/pa.float32()` reproduces the same failure, as does calling `entries()`, `data_files()` or `all_files()` instead of `files()`.

### Expected behavior

`inspect.files()` and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.

### Analysis

`InspectTable._get_files_from_manifest` decodes `lower_bounds` / `upper_bounds` via `from_bytes(field.field_type, ...)`, i.e. using the **current** schema type.

Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after `int` → `long`, pre-existing files still carry 4-byte bounds while the current field type is `long`, and `_LONG_STRUCT.unpack` (8 bytes) fails.

As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the **encoded byte length** rather than trusting the current schema type.

### Workaround

`Table.scan().plan_files()` returns `DataFile` objects without decoding bounds, so file-level metadata (`file_path`, `record_count`, `sort_order_id`, `spec_id`, …) remains reachable:

```python
tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
```

### Environment

- pyiceberg 0.11.1 (latest release)
- pyarrow 25.0.0, SQLAlchemy 2.0.51
- Python 3.14.5, macOS (arm64)

Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose `int` column had been promoted to `bigint` via Athena; the reproduction above shows it is catalog-independent.

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Tái hiện lỗi bằng script tự chứa, sau đó đọc pyiceberg/table/inspect.py tại InspectTable._get_files_from_manifest và pyiceberg/conversions.py tại from_bytes. So sánh độ dài bound đã mã hóa với các kiểu trường hiện tại cho int→long và float→double; được xem là hoàn tất khi files(), entries(), data_files() và all_files() trả về bound hoặc None mà không phát sinh struct.error.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python
Lĩnh vực
databases
Loại issue
Lỗi
Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức độ hoạt động
Ít trao đổi
Độ rõ ràng
Đặc tả rõ ràng
Mức phù hợp với người mới
68/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.