apache / apache/iceberg-python
Pluggable Backend Interface with DataFusion for Bounded-Memory Compute
- 主要言語
- Python
- スター
- 1.1k
- フォーク
- 581
- 平均マージ
- 1日 17時間
- マージ済み PR(30日)
- 78
説明
## Summary
PyIceberg uses PyArrow as its sole execution engine. PyArrow is a kernel library with no memory management, no spill-to-disk, and no join operators. Operations that process more data than available memory (CoW deletes, equality delete resolution, scan planning for heavily-deleted tables, sorted writes) crash with OOM errors.
This issue tracks introducing a pluggable backend interface (`ReadBackend`, `WriteBackend`, `ComputeBackend` protocols) and integrating Apache DataFusion as the first bounded-memory compute backend.
## Problem
| Operation | Current Status | OOM Pattern |
|-----------|---------------|-------------|
| Equality delete reads | Hard `ValueError` | Anti-join requires all delete keys in memory |
| CoW delete (large files) | OOMs | Materializes entire Parquet file into RAM |
| Scan planning (>100K deletes) | OOMs | All delete entries in Python dict |
| Sort-on-write | Not implemented | Full sort before write |
| Positional deletes (millions) | OOMs | Python set of positions |
Tables written by Flink (which uses equality deletes) are completely unreadable by PyIceberg today.
## Solution
1. **Pluggable interface**: `ReadBackend`, `WriteBackend`, `ComputeBackend` protocols that decouple PyIceberg from PyArrow
2. **DataFusion integration**: Bounded-memory sort, join, and filter with spill-to-disk via `datafusion-python`
3. **Migration**: All existing data operations route through the interface with zero API changes
## Deliverables
- [ ] Equality delete resolution (NEW): tables with equality deletes can now be read
- [ ] CoW delete/overwrite streaming (FIX): statistics short-circuit + two-pass streaming
- [ ] Positional delete resolution (IMPROVED): bounded-memory for large delete sets
- [ ] Sort-on-write (NEW): external merge sort when DataFusion installed
- [ ] Bounded-memory scan planning (NEW): for tables with >100K delete files
## Related Issues
- #1210 - Support reading equality delete files
- #3270 - Equality Delete support
- #3554 - Integrate DataFusion as execution engine
## Acceptance Criteria
- All existing tests pass without `datafusion` installed (no regression)
- Tables with equality deletes return correct results
- CoW delete on 2GB+ files completes without OOM (with DataFusion)
- Sort-on-write produces sorted files when table has sort order and DataFusion installed
- Property-based tests verify PyArrow and DataFusion backends produce identical output
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
調査の方向性
実装ファイルやテストは指定されていません。まず関連する issue #1210、#3270、#3554 を読み、次に提案されている ReadBackend、WriteBackend、ComputeBackend プロトコルと DataFusion 統合の範囲を定めてください。完了条件は、一覧にある delete、ストリーミング、ソート、スキャン計画の動作が実現され、DataFusion なしでもリグレッションがなく、PyArrow と DataFusion の結果が同等であることです。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- backend, data-engineering, databases
- issue の種類
- 機能追加
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 静か
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 25/100