apache / apache/parquet-java

Improve logic when to write column indexes

Aperta
#2,228 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Component: Parquet Priority: Minor Type: enhancement
Lingua principale
Java
Stelle
3.1k
Fork
1.6k
Merge medio
3g 12h
PR unite (30g)
33

Descrizione

Currently, we always write column indexes. In case of the data is ordered (ASCENDING or DESCENDING) the filtering would highly benefit from column indexes. While, if the data is UNORDERED it is not obvious if ordering based on column indexes would make sense. For example if the data is random then the min/max values of the different pages might be close to each other so in most cases filtering based on these values would not drop any of the pages. In the other hand UNORDERED values does not mean that the values are random. It can happen that the values are clustered or semi-ordered. We shall discover these cases somehow before writing the column indexes and write only if the min/max values for the pages do not overlap too much.

Another simple case if we have only one page. In this case writing column indexes is useless. 

**Reporter**: [Gabor Szadovszky](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=gszadovszky) / @gszadovszky
**Assignee**: [Gabor Szadovszky](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=gszadovszky) / @gszadovszky
#### Related issues:
- [Column indexes](https://github.com/apache/parquet-java/issues/2123) (depends upon)
- [Benchmark filtering column-indexes](https://github.com/apache/parquet-java/issues/2235) (depends upon)

**Note**: *This issue was originally created as [PARQUET-1415](https://issues.apache.org/jira/browse/PARQUET-1415). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

Inizia tracciando il percorso di scrittura di column-index e il modo in cui gestisce i dati ordinati rispetto a quelli non ordinati. Confronta il comportamento per i dati su una singola pagina e in caso di sovrapposizione dei valori min/max delle pagine; il lavoro è completato quando gli indexes vengono omessi se non possono migliorare il filtraggio, con copertura dei casi ordinati, non ordinati, raggruppati e a pagina singola.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java
Ambito
data-engineering, performance
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.