apache / apache/parquet-java

ColumnIndex should provide number of records skipped

Aperta
#2,536 17 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Component: Java Component: Parquet Priority: Major Type: enhancement
Lingua principale
Java
Stelle
3.1k
Fork
1.6k
Merge medio
3g 12h
PR unite (30g)
33

Descrizione

When integrating Parquet ColumnIndex, I found we need to know from Parquet that how many records that we skipped due to ColumnIndex filtering. When rowCount is 0, readNextFilteredRowGroup() just advance to next without telling the caller. See code here

 

In Iceberg, it reads Parquet record with an iterator. The hasNext() has the following code():

valuesRead + skippedValues < totalValues

See ( 

So without knowing the skipped values, it is hard to determine hasNext() or not. 

 

Currently, we can workaround by using a flag. When readNextFilteredRowGroup() returns null, we consider it is done for the whole file. Then hasNext() just retrun false. 

 

 

 

**Reporter**: [Xinli Shang](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=shangxinli) / @shangxinli
**Assignee**: [Xinli Shang](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=shangx@uber.com) / @shangxinli

**Note**: *This issue was originally created as [PARQUET-1927](https://issues.apache.org/jira/browse/PARQUET-1927). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

Inizia da parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java, in readNextFilteredRowGroup(), quindi confronta la logica hasNext() dell’iteratore nella modifica Iceberg collegata. Traccia il modo in cui avanza un row group con rowCount 0 e determina come il conteggio dei record ignorati debba arrivare al chiamante; il lavoro è completo quando i chiamanti possono determinare in modo affidabile se rimangono record dopo il filtraggio di ColumnIndex.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java
Ambito
data-engineering
Tipo di issue
Funzionalità
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.