apache / apache/iceberg-python

Metadata inspection APIs fail with struct.error after int→long / float→double type promotion

Offen
#3,744 3 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
1.1k
Forks
581
Ø Merge
1 T. 17 Std.
Gemergte PRs (30 T.)
78

Beschreibung

### Apache Iceberg version

0.11.0 (latest release)

### Please describe the bug 🐞

### Summary

After a type promotion that the spec allows (`int` → `long`, `float` → `double`), **all metadata inspection APIs raise `struct.error`** on any table that already contains data files written before the promotion:

| Promotion | `inspect.files()` | `inspect.entries()` | `inspect.data_files()` | `inspect.all_files()` |
|---|---|---|---|---|
| `int` → `long` | ❌ | ❌ | ❌ | ❌ |
| `float` → `double` | ❌ | ❌ | ❌ | ❌ |

```
struct.error: unpack requires a buffer of 8 bytes
```

Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.

### Reproduction

Self-contained, no cloud services or network required:

```python
import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType

warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")

tbl = catalog.create_table("ns.t", schema=Schema(
NestedField(1, "name", StringType(), required=False),
NestedField(2, "qty", IntegerType(), required=False),
))

# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
[{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
schema=pa.schema([pa.field("name", pa.string(), nullable=True),
pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")

# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
update.update_column("qty", field_type=LongType())

tbl = catalog.load_table("ns.t")
try:
tbl.inspect.files()
except Exception:
traceback.print_exc()

shutil.rmtree(warehouse, ignore_errors=True)
```

**Output**

```
before promotion: 1 row(s) -> OK
Traceback (most recent call last):
...
File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
"lower_bound": from_bytes(field.field_type, lower_bound)
File ".../pyiceberg/conversions.py", line 337, in _
return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes
```

Replacing `IntegerType()/LongType()/pa.int32()` with `FloatType()/DoubleType()/pa.float32()` reproduces the same failure, as does calling `entries()`, `data_files()` or `all_files()` instead of `files()`.

### Expected behavior

`inspect.files()` and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.

### Analysis

`InspectTable._get_files_from_manifest` decodes `lower_bounds` / `upper_bounds` via `from_bytes(field.field_type, ...)`, i.e. using the **current** schema type.

Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after `int` → `long`, pre-existing files still carry 4-byte bounds while the current field type is `long`, and `_LONG_STRUCT.unpack` (8 bytes) fails.

As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the **encoded byte length** rather than trusting the current schema type.

### Workaround

`Table.scan().plan_files()` returns `DataFile` objects without decoding bounds, so file-level metadata (`file_path`, `record_count`, `sort_order_id`, `spec_id`, 
) remains reachable:

```python
tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
```

### Environment

- pyiceberg 0.11.1 (latest release)
- pyarrow 25.0.0, SQLAlchemy 2.0.51
- Python 3.14.5, macOS (arm64)

Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose `int` column had been promoted to `bigint` via Athena; the reproduction above shows it is catalog-independent.

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Beitragsleitfaden

FĂŒr dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Reproduce the failure with the self-contained script, then read pyiceberg/table/inspect.py at InspectTable._get_files_from_manifest and pyiceberg/conversions.py at from_bytes. Compare the encoded bound lengths with the current field types for int→long and float→double; done means files(), entries(), data_files(), and all_files() return bounds or None without raising struct.error.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
databases
Issue-Typ
Bug
Schwierigkeit
3/5
GeschÀtzter Aufwand
1-2 Tage
AktivitÀtsstatus
Ruhig
Klarheit
Klar beschrieben
AnfÀngerfreundlichkeit
68/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht ĂŒber anfĂ€ngerfreundliche GitHub-Issues.