apache / apache/iceberg-python

Honor identity sort orders for Arrow table writes

Aberta
#3,848 0 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
1.1k
Forks
581
Merge médio
1d 13h
PRs com merge (30d)
76

Descrição

## Feature Request / Improvement

PyIceberg accepts table sort orders at write time but ignores them. Arrow table writes produce unsorted data files and hard-code `sort_order_id=None`, so manifests are inconsistent with the sort order the user declared.

Related umbrella issue: #271

### Use case / motivation

Without a truthful sort_order_id and physically ordered data files, readers can't use sort-order-aware pruning and the manifest disagrees with the table's sort metadata. Users who declare a sort order expect their writes to honor it.

### Proposed change

When every sort field uses an identity transform and null placement is consistent, sort materialized Arrow table writes. Unpartitioned tables are sorted before bin packing and each final partition is sorted independently. The table's sort-order ID is recorded on the data files.

Unsupported transforms, nested or missing fields, mixed null placement, and streaming RecordBatchReader writes keep the current behavior. Files are marked unsorted and sort_order_id stays null, with a warning explaining why.

### Implementation

PR #3830 sorts the writes, carries the sort-order ID through WriteTask, writes the truthful DataFile.sort_order_id, and includes unit and integration coverage.

### Tooling note

I developed this with assistance from DS v4 Pro and reviewed the changes myself.

### References

- #271
- #3830

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Direção de pesquisa

Revise primeiro o PR #3830, juntamente com a cobertura de testes unitários e de integração que ele adiciona para gravações de tabelas Arrow. Verifique como sort-order IDs, WriteTask, DataFile, manifests e bin packing são tratados. Está concluído quando as identity-sort writes suportadas produzem arquivos ordenados com valores sort_order_id verdadeiros, enquanto os casos não suportados mantêm o comportamento e o aviso atuais.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
python
Domínio
databases
Tipo de issue
Funcionalidade
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Estagnada
Clareza
Razoavelmente clara
Facilidade para iniciantes
25/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.