apache / apache/iceberg-python
Metadata inspection APIs fail with struct.error after int→long / float→double type promotion
- Dominant language
- Python
- Stars
- 1.1k
- Forks
- 581
- Avg merge
- 1d 17h
- Merged PRs (30d)
- 78
Description
### Apache Iceberg version
0.11.0 (latest release)
### Please describe the bug 🐞
### Summary
After a type promotion that the spec allows (`int` → `long`, `float` → `double`), **all metadata inspection APIs raise `struct.error`** on any table that already contains data files written before the promotion:
| Promotion | `inspect.files()` | `inspect.entries()` | `inspect.data_files()` | `inspect.all_files()` |
|---|---|---|---|---|
| `int` → `long` | ❌ | ❌ | ❌ | ❌ |
| `float` → `double` | ❌ | ❌ | ❌ | ❌ |
```
struct.error: unpack requires a buffer of 8 bytes
```
Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.
### Reproduction
Self-contained, no cloud services or network required:
```python
import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType
warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")
tbl = catalog.create_table("ns.t", schema=Schema(
NestedField(1, "name", StringType(), required=False),
NestedField(2, "qty", IntegerType(), required=False),
))
# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
[{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
schema=pa.schema([pa.field("name", pa.string(), nullable=True),
pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")
# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
update.update_column("qty", field_type=LongType())
tbl = catalog.load_table("ns.t")
try:
tbl.inspect.files()
except Exception:
traceback.print_exc()
shutil.rmtree(warehouse, ignore_errors=True)
```
**Output**
```
before promotion: 1 row(s) -> OK
Traceback (most recent call last):
...
File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
"lower_bound": from_bytes(field.field_type, lower_bound)
File ".../pyiceberg/conversions.py", line 337, in _
return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes
```
Replacing `IntegerType()/LongType()/pa.int32()` with `FloatType()/DoubleType()/pa.float32()` reproduces the same failure, as does calling `entries()`, `data_files()` or `all_files()` instead of `files()`.
### Expected behavior
`inspect.files()` and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.
### Analysis
`InspectTable._get_files_from_manifest` decodes `lower_bounds` / `upper_bounds` via `from_bytes(field.field_type, ...)`, i.e. using the **current** schema type.
Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after `int` → `long`, pre-existing files still carry 4-byte bounds while the current field type is `long`, and `_LONG_STRUCT.unpack` (8 bytes) fails.
As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the **encoded byte length** rather than trusting the current schema type.
### Workaround
`Table.scan().plan_files()` returns `DataFile` objects without decoding bounds, so file-level metadata (`file_path`, `record_count`, `sort_order_id`, `spec_id`, …) remains reachable:
```python
tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
```
### Environment
- pyiceberg 0.11.1 (latest release)
- pyarrow 25.0.0, SQLAlchemy 2.0.51
- Python 3.14.5, macOS (arm64)
Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose `int` column had been promoted to `bigint` via Athena; the reproduction above shows it is catalog-independent.
### Willingness to contribute
- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time
Contributor guide
No contributing guide indexed for this repository
Research direction
Reproduce the failure with the self-contained script, then read pyiceberg/table/inspect.py at InspectTable._get_files_from_manifest and pyiceberg/conversions.py at from_bytes. Compare the encoded bound lengths with the current field types for int→long and float→double; done means files(), entries(), data_files(), and all_files() return bounds or None without raising struct.error.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- databases
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Quiet
- Clarity
- Clearly specified
- Newbie friendliness
- 68/100