apache / apache/parquet-java

FilteredRecordReader skips rows it shouldn't for schema with optional columns

Ouverte
#1,730 6 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Component: Java Component: Parquet Priority: Blocker Type: bug
Langage dominant
Java
Étoiles
3.1k
Forks
1.6k
Merge moyen
3 j 12 h
PR mergées (30 j)
33

Description

When using UnboundRecordFilter with nested AND/OR filters over OPTIONAL columns, there seems to be a case with a mismatch between the current record's column value and the value read during filtering.

The structure of my filter predicate that results in incorrect filtering is: (x && (y || z))

When I step through it with a debugger I can see that the value being read from the ColumnReader inside my Predicate is different than the value for that row.

Looking deeper there seems to be a buffer with dictionary keys in RunLenghBitPackingHybridDecoder (I am using RLE). There are only two different keys in this array, [0,1], whereas my optional column has three different values, [null,0,1]. If I had a column with values 5,10,10,null,10, and keys 0 -> 5 and 1 -> 10, the buffer would hold 0,1,1,1,0, and in the case that it reads the last row, would return 0 -> 5.

So it seems that nothing is keeping track of where nulls appear.

Hope someone can take a look, as it is a blocker for my project.

**Environment**: Linux, Java7/Java8
**Reporter**: [Steven Mellinger](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=stevemel)
#### Related issues:
- [filter2 API performance regression](https://github.com/apache/parquet-java/issues/1583) (relates to)

**Note**: *This issue was originally created as [PARQUET-182](https://issues.apache.org/jira/browse/PARQUET-182). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Commencez par suivre le prédicat imbriqué (x \u0026\u0026 (y || z)) à travers FilteredRecordReader, UnboundRecordFilter et ColumnReader. Inspectez RunLenghBitPackingHybridDecoder avec une colonne facultative contenant null, 0 et 1, puis comparez les valeurs décodées aux positions des lignes. C'est terminé lorsque le filtrage renvoie les bonnes lignes et que la couverture de régression préserve les positions null.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
java
Domaine
data-engineering
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.