apache / apache/iceberg-python
Upsert with 1M rows extremely slow due to `create_match_filter` and `txn.delete()` performance
- Linguagem predominante
- Python
- Estrelas
- 1.1k
- Forks
- 581
- Merge médio
- 1d 17h
- PRs com merge (30d)
- 78
Descrição
### Apache Iceberg version
0.11.0
### Please describe the bug 🐞
CC @goutamvenkat-anyscale @koenvo @Fokko
Hello! We are implementing distributed writes from Ray Data to Iceberg. As part of upserts, we:
1. Write data files in parallel across Ray workers (each worker writes its share of Parquet files directly to storage and returns `DataFile` metadata + the upsert key columns back to the driver)
2. On the driver, concatenate all upsert keys collected from workers, call `create_match_filter` to build a delete predicate, then call `txn.delete()` followed by an append to commit
Upserting 1M rows (383 MiB) into an Iceberg table takes **~17.5 minutes**, almost entirely in the delete step:
```
create_match_filter (1M keys → In filter): 10.26s
txn.delete(): 1054.35s
append + commit: 1.14s
─────────────────────────────────────────────────────
Total upsert commit: 1065.75s
```
PyIceberg version `0.11.0`
This matches what's reported in #2159 and #2138.
The bottlenecks are:
1. **`create_match_filter`** — constructs a Python `BooleanExpression` node per row, which is expensive at 1M+ keys
2. **`txn.delete()`** — evaluates the resulting giant `In` expression against the table's data files with no partition pruning, effectively doing a full table scan
We have a few questions:
1. **Merge-on-read upserts** — is this on the roadmap, and if so, roughly when? MoR would let us avoid the expensive delete + rewrite cycle entirely for large upserts.
2. **Optimizing `create_match_filter` or `txn.delete()`** — is there a recommended way to speed these up today? For example, batching the `In` filter, or passing a partition-level hint to constrain the file scan?
3. **Partition-aware deletes** — if the upsert key columns overlap with partition columns, is there a supported way to restrict `txn.delete()` to only the relevant partitions, rather than scanning the full table?
## Related
- #2159 — Upserting large table extremely slow
- #2138 — Upsertion memory usage grows exponentially as table size grows
- #2943 — Optimize upsert performance for large datasets
### Willingness to contribute
- [ ] I can contribute a fix for this bug independently
- [x] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time
Guia de contribuição
Nenhum guia de contribuição indexado para este repositório
Direção de pesquisa
Comece pelos pontos de entrada create_match_filter e txn.delete() descritos no relatório e, em seguida, reproduza o benchmark de upsert de 1M de linhas com base nos tempos fornecidos. Consulte as issues relacionadas #2159, #2138 e #2943 para obter o contexto existente. O trabalho estará concluído quando houver uma melhoria medida ou uma forma documentada e compatível de evitar o custo relatado da exclusão da tabela inteira.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Stack de tecnologia
- python
- Domínio
- data-engineering, databases
- Tipo de issue
- Bug
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Status de atividade
- Pouca atividade
- Clareza
- Precisa de esclarecimento
- Facilidade para iniciantes
- 35/100