apache / apache/parquet-java

toParquetMetadata method in ParquetMetadataConverter does not set dictionary page offset bit

Aperta
#2,901 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Component: Java Component: Parquet Priority: Major Type: bug
Lingua principale
Java
Stelle
3.1k
Fork
1.6k
Merge medio
3g 12h
PR unite (30g)
33

Descrizione

toParquetMetadata method converts org.apache.parquet.hadoop.metadata.ParquetMetadata to org.apache.parquet.format.FileMetaData but this does not set the dictionary page offset bit in FileMetaData.

When a FileMetaData object is serialized while writing to the footer and then deserialized, the dictionary offset is lost as the dictionary page offset bit was never set.

PARQUET-1850  tried to fix this but it did only a partial fix.

It sets setDictionary_page_offset only if getEncodingStats are present
```java

if (columnMetaData.getEncodingStats() != null
&& columnMetaData.getEncodingStats().hasDictionaryPages())
{ metaData.setDictionary_page_offset(columnMetaData.getDictionaryPageOffset()); }
```
However, it should setDictionary_page_offset even when getEncodingStats are not present but encodings are present.

It should use the implementation in ColumnChunkMetatdata below:
```java

public boolean hasDictionaryPage() {
EncodingStats stats = getEncodingStats();
if (stats != null) {
return stats.hasDictionaryPages() && stats.hasDictionaryEncodedPages();
}

Set encodings = getEncodings();
return (encodings.contains(PLAIN_DICTIONARY) || encodings.contains(RLE_DICTIONARY));
}
```
So new change in ParquetMetadataCOnvertor should be like:

 
```java

if (columnMetaData.hasDictionaryPage()) { metaData.setDictionary_page_offset(columnMetaData.getDictionaryPageOffset()); }
```

**Reporter**: [Abhishek Dixit](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=abhishekd0907)
#### PRs and other links:
- [GitHub Pull Request #1340](https://github.com/apache/parquet-mr/pull/1340)

**Note**: *This issue was originally created as [PARQUET-2464](https://issues.apache.org/jira/browse/PARQUET-2464). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

Inizia da ParquetMetadataConverter.toParquetMetadata e confronta la sua condizione per le pagine del dizionario con ColumnChunkMetadata.hasDictionaryPage(), come descritto nell’issue. Verifica che la serializzazione e la deserializzazione del footer preservino l’offset della pagina del dizionario quando le statistiche di encoding sono assenti ma sono presenti gli encoding del dizionario; controlla Pull Request #1340 per verificare il lavoro già esistente.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java
Ambito
data-engineering
Tipo di issue
Bug
Difficoltà
2/5
Tempo stimato
1-3 ore
Stato di attività
Ferma
Chiarezza
Specificata chiaramente
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.