perf: push_down_filter is pathologically slow for some plans
- Lingua principale
- Rust
- Stelle
- 9.3k
- Fork
- 2.4k
- Merge medio
- 3g 11h
- PR unite (30g)
- 362
Descrizione
### Describe the bug
While investigating #17261 it became apparent that one of the largest consumers of cpu time during planning of the sql_planner_extended benchmark is the PushDownFilter OptimizerRule. I've instrumented datafusion with some logging and ran the benchmark with the following cmd:
`RUST_LOG=info cargo samply --profile=release-nonlto --bench sql_planner_extended -- --nocapture --sample-size 10`
The full output can be seen in [this gist](https://gist.github.com/Omega359/978e208b401f6af03fdf00fd8af63938) but below is the pertinent bit:
```
[2026-01-25T15:40:20Z INFO datafusion_optimizer::optimizer] Optimization (round 0) for rule push_down_limit took > 50ms: 90ms
[2026-01-25T15:43:14Z INFO datafusion_optimizer::optimizer] Optimization (round 0) for rule push_down_filter took > 50ms: 174174ms
[2026-01-25T15:43:15Z INFO datafusion_optimizer::optimizer] Optimization (round 0) for rule single_distinct_aggregation_to_group_by took > 50ms: 164ms
[2026-01-25T15:43:15Z INFO datafusion_optimizer::optimizer] Optimization (round 0) for rule eliminate_group_by_constant took > 50ms: 159ms
[2026-01-25T15:43:16Z INFO datafusion_optimizer::optimizer] Optimization (round 0) for rule common_sub_expression_eliminate took > 50ms: 1313ms
[2026-01-25T15:43:20Z INFO datafusion_optimizer::optimizer] Optimization (round 0) for rule optimize_projections took > 50ms: 3389ms
```
As you can see quite a few optimizer rules are using too much cpu for planning however the push_down_filter is the most egregious taking 174 seconds to complete. You can see from a screenshot of the output of samply where it seems most of that time is going.
### To Reproduce
`RUST_LOG=info cargo samply --profile=release-nonlto --bench sql_planner_extended -- --nocapture --sample-size 10`
Branch with profiling log @ https://github.com/Omega359/arrow-datafusion/tree/profile_optimize
### Expected behavior
Plan optimization should not be exponentially slow for some logical plans.
### Additional context
_No response_
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia eseguendo il benchmark sql_planner_extended con il comando cargo samply fornito ed esaminando l’output del profiling. Ispeziona la OptimizerRule PushDownFilter e il branch profile_optimize per individuare l’hotspot di pianificazione segnalato. Il lavoro è completato quando il benchmark non mostra più un tempo di pianificazione patologico per i piani logici interessati.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- rust
- Ambito
- performance
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Ferma
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 35/100