apache / apache/iceberg-python

IO: Consolidate PyArrow logic into io/pyarrow.py before decomposition

Offen
#3,812 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
1.1k
Forks
581
Ø Merge
1 T. 17 Std.
Gemergte PRs (30 T.)
77

Beschreibung

## Summary

Before decomposing `pyiceberg/io/pyarrow.py` into focused submodules (#3737, #3738), we should consolidate PyArrow-specific logic that currently lives outside the module. This ensures all PyArrow calls route through a single boundary, making the subsequent split clean and enabling future engine substitution.

## Motivation

Per discussion in #3737, @rambleraptor noted that the first useful step is ensuring no PyArrow logic occurs outside `pyarrow.py`. Currently several modules import `pyarrow` directly and implement compute logic inline rather than delegating through `pyiceberg.io.pyarrow`.

When we later introduce a `ComputeEngine` protocol, any PyArrow logic outside the module boundary bypasses the protocol and prevents clean substitution.

## Audit

Grepped `pyiceberg/` (excluding `io/pyarrow.py` and tests) for runtime `import pyarrow` statements (both top-level and inline). Excluded `TYPE_CHECKING`-only imports since those have no runtime dependency.

| Location | What it does | Action |
|----------|-------------|--------|
| `table/upsert_util.py` | PyArrow table joins, group_by, compute, cast, take | **Absorb** |
| `table/inspect.py` | Builds pa.schema + pa.Table.from_pylist for metadata inspection | **TBD** |
| `transforms.py` | `pyarrow_transform()` dispatch on pa.Array/ChunkedArray | **TBD** |
| `table/__init__.py` | Entry points accept pa.Table, delegate to io.pyarrow | **Leave** |
| `table/deletion_vector.py` | Single pa.chunked_array() call | **Leave** |
| `catalog/__init__.py` | Delegates to io.pyarrow for schema conversion | **Leave** |

## Plan

One PR per absorption. Each is a pure refactor: move code into `io/pyarrow.py`, have the caller import from `pyiceberg.io.pyarrow` instead of `pyarrow` directly. No behavior change, all existing tests pass unchanged.

- [ ] PR A: Absorb `table/upsert_util.py` PyArrow logic
- [ ] PR B: `table/inspect.py` (pending discussion)
- [ ] PR C: `transforms.py` (pending discussion)

## Related

- #3737 - Decompose io/pyarrow.py into focused modules
- #3738 - Extract PyArrowFileIO (first decomposition step)
- #3715 / #3716 - Previous pluggable backend attempt (rejected as too large)

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Beginne damit, table/upsert_util.py mit pyiceberg/io/pyarrow.py zu vergleichen und die zugehörige Diskussion in #3737 zu prüfen. Konzentriere dich zuerst auf PR A; abgeschlossen bedeutet, dass die PyArrow-Logik des Upsert-Utilities über io/pyarrow.py läuft, ohne Verhaltensänderungen und mit erfolgreich durchlaufenden bestehenden Tests.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
backend
Issue-Typ
Refactoring
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Aktiv
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
55/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.