Investigate performance tradeoff in compressing spill files
- Lingua principale
- Rust
- Stelle
- 9.3k
- Fork
- 2.4k
- Merge medio
- 3g 11h
- PR unite (30g)
- 360
Descrizione
### Is your feature request related to a problem or challenge?
Part of https://github.com/apache/datafusion/issues/16065 , Related to https://github.com/apache/datafusion/issues/14078
### Background
PR https://github.com/apache/datafusion/pull/16268 will introduce compression option when writing spill files in disk. These options are limited to compression options to what Arrow IPC Stream Writer supports : `zstd`, `lz4_frame` (compression level is not configurable)
To further enhance spilling execution, we need to investigate
- CPU / I/O tradeoff when `zstd` or `lz4_frame` compression is enabled i.e. compression ratio, extra latency spent for compression
- Current arrow ipc stream writer always write `batch` at a time in `append_batch`. In terms of compression, it is not sure yet how much single batch can benefit from compression.
- whether we need separate `Writer` or `Reader` implementation instead of IPC Stream Writer.
- how to introduce sort of `adaptiveness`.
### Describe the solution you'd like
First, we need to track (or update) how many bytes are written in spill files. Datafusion currently tracks `spilled_bytes` as part of `SpillMetrics`, but it is calculated based on in memory array size, which would be different from actual spill files size especially when we compress spill files.
Second, update the benchmarks or write a separate benchmarks to see the performance characteristics. One possible way is writing out spill-related metrics to output.json when running benches like tpch with `debug` option. Another idea is to generate some spill files for microbenchmark testing only spill writing - reading process.
### Describe alternatives you've considered
_No response_
### Additional context
_No response_
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia tracciando SpillMetrics e il calcolo di spilled_bytes, quindi esamina il percorso di spill di Arrow IPC Stream Writer e i benchmark tpch esistenti con l’opzione debug. Confronta la scrittura e la lettura di spill compressi e non compressi, incluso il rapporto di compressione e la latenza. Il lavoro è completato quando vengono tracciati i byte effettivi dei file di spill e l’output del benchmark espone metriche sufficienti per valutare i compromessi.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- rust
- Ambito
- data-engineering, performance
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Ferma
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100