apache / apache/iceberg-python

Honor identity sort orders for Arrow table writes

Abierto
#3,848 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
1.1k
Forks
581
Merge medio
1 d 17 h
PR fusionados (30 d)
77

Descripción

## Feature Request / Improvement

PyIceberg accepts table sort orders at write time but ignores them. Arrow table writes produce unsorted data files and hard-code `sort_order_id=None`, so manifests are inconsistent with the sort order the user declared.

Related umbrella issue: #271

### Use case / motivation

Without a truthful sort_order_id and physically ordered data files, readers can't use sort-order-aware pruning and the manifest disagrees with the table's sort metadata. Users who declare a sort order expect their writes to honor it.

### Proposed change

When every sort field uses an identity transform and null placement is consistent, sort materialized Arrow table writes. Unpartitioned tables are sorted before bin packing and each final partition is sorted independently. The table's sort-order ID is recorded on the data files.

Unsupported transforms, nested or missing fields, mixed null placement, and streaming RecordBatchReader writes keep the current behavior. Files are marked unsorted and sort_order_id stays null, with a warning explaining why.

### Implementation

PR #3830 sorts the writes, carries the sort-order ID through WriteTask, writes the truthful DataFile.sort_order_id, and includes unit and integration coverage.

### Tooling note

I developed this with assistance from DS v4 Pro and reviewed the changes myself.

### References

- #271
- #3830

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Revisa primero el PR #3830, junto con la cobertura de unit e integración que añade para las escrituras de tablas Arrow. Comprueba cómo se gestionan los sort-order IDs, WriteTask, DataFile, manifests y bin packing. Se considera completado cuando las identity-sort writes compatibles producen archivos ordenados con valores de sort_order_id veraces, mientras que los casos no compatibles conservan el comportamiento y la advertencia actuales.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python
Área
databases
Tipo de issue
Nueva funcionalidad
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.