apache / apache/iceberg-python

Metadata inspection APIs fail with struct.error after int→long / float→double type promotion

Ouverte
#3,744 3 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Python
Étoiles
1.1k
Forks
581
Merge moyen
1 j 17 h
PR mergées (30 j)
77

Description

### Apache Iceberg version

0.11.0 (latest release)

### Please describe the bug 🐞

### Summary

After a type promotion that the spec allows (`int` → `long`, `float` → `double`), **all metadata inspection APIs raise `struct.error`** on any table that already contains data files written before the promotion:

| Promotion | `inspect.files()` | `inspect.entries()` | `inspect.data_files()` | `inspect.all_files()` |
|---|---|---|---|---|
| `int` → `long` | ❌ | ❌ | ❌ | ❌ |
| `float` → `double` | ❌ | ❌ | ❌ | ❌ |

```
struct.error: unpack requires a buffer of 8 bytes
```

Since type promotion is not rewriting existing data files, the table stays in this state permanently — the whole metadata-inspection surface becomes unusable.

### Reproduction

Self-contained, no cloud services or network required:

```python
import tempfile, shutil, traceback
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.schema import Schema
from pyiceberg.types import NestedField, IntegerType, LongType, StringType

warehouse = tempfile.mkdtemp(prefix="iceberg_repro_")
catalog = SqlCatalog("repro", uri=f"sqlite:///{warehouse}/catalog.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")

tbl = catalog.create_table("ns.t", schema=Schema(
NestedField(1, "name", StringType(), required=False),
NestedField(2, "qty", IntegerType(), required=False),
))

# Write while the column is still `int` -> bounds are stored as 4-byte LE.
tbl.append(pa.Table.from_pylist(
[{"name": "a", "qty": 1}, {"name": "b", "qty": 2}],
schema=pa.schema([pa.field("name", pa.string(), nullable=True),
pa.field("qty", pa.int32(), nullable=True)]),
))
print("before promotion:", tbl.inspect.files().num_rows, "row(s) -> OK")

# Allowed promotion; existing data files and bounds are not rewritten.
with tbl.update_schema() as update:
update.update_column("qty", field_type=LongType())

tbl = catalog.load_table("ns.t")
try:
tbl.inspect.files()
except Exception:
traceback.print_exc()

shutil.rmtree(warehouse, ignore_errors=True)
```

**Output**

```
before promotion: 1 row(s) -> OK
Traceback (most recent call last):
...
File ".../pyiceberg/table/inspect.py", line 573, in _get_files_from_manifest
"lower_bound": from_bytes(field.field_type, lower_bound)
File ".../pyiceberg/conversions.py", line 337, in _
return _LONG_STRUCT.unpack(b)[0]
struct.error: unpack requires a buffer of 8 bytes
```

Replacing `IntegerType()/LongType()/pa.int32()` with `FloatType()/DoubleType()/pa.float32()` reproduces the same failure, as does calling `entries()`, `data_files()` or `all_files()` instead of `files()`.

### Expected behavior

`inspect.files()` and the other metadata tables should return the bounds (or omit/None them) rather than raising, on tables that have undergone a spec-allowed type promotion.

### Analysis

`InspectTable._get_files_from_manifest` decodes `lower_bounds` / `upper_bounds` via `from_bytes(field.field_type, ...)`, i.e. using the **current** schema type.

Per the spec, type promotion does not rewrite existing bounds, and the Avro manifest does not record which type was used to encode them. So after `int` → `long`, pre-existing files still carry 4-byte bounds while the current field type is `long`, and `_LONG_STRUCT.unpack` (8 bytes) fails.

As discussed on the dev list regarding type promotion in v3, implementations handle this by detecting the promotion from the **encoded byte length** rather than trusting the current schema type.

### Workaround

`Table.scan().plan_files()` returns `DataFile` objects without decoding bounds, so file-level metadata (`file_path`, `record_count`, `sort_order_id`, `spec_id`, …) remains reachable:

```python
tasks = list(tbl.scan().plan_files())
[(t.file.file_path, t.file.record_count) for t in tasks]
```

### Environment

- pyiceberg 0.11.1 (latest release)
- pyarrow 25.0.0, SQLAlchemy 2.0.51
- Python 3.14.5, macOS (arm64)

Originally hit against an AWS S3 Tables table (Glue Iceberg REST catalog) whose `int` column had been promoted to `bigint` via Athena; the reproduction above shows it is catalog-independent.

### Willingness to contribute

- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Reproduisez l’échec avec le script autonome, puis consultez pyiceberg/table/inspect.py à InspectTable._get_files_from_manifest et pyiceberg/conversions.py à from_bytes. Comparez les longueurs des bounds encodées avec les types de champs actuels pour int→long et float→double ; le travail est terminé lorsque files(), entries(), data_files() et all_files() renvoient des bounds ou None sans lever struct.error.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
databases
Type d'issue
Bug
Difficulté
3/5
Temps estimé
1-2 jours
Activité
Calme
Clarté
Clairement spécifiée
Accessibilité débutants
68/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.