FilteredRecordReader skips rows it shouldn't for schema with optional columns
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 33
Description
When using UnboundRecordFilter with nested AND/OR filters over OPTIONAL columns, there seems to be a case with a mismatch between the current record's column value and the value read during filtering.
The structure of my filter predicate that results in incorrect filtering is: (x && (y || z))
When I step through it with a debugger I can see that the value being read from the ColumnReader inside my Predicate is different than the value for that row.
Looking deeper there seems to be a buffer with dictionary keys in RunLenghBitPackingHybridDecoder (I am using RLE). There are only two different keys in this array, [0,1], whereas my optional column has three different values, [null,0,1]. If I had a column with values 5,10,10,null,10, and keys 0 -> 5 and 1 -> 10, the buffer would hold 0,1,1,1,0, and in the case that it reads the last row, would return 0 -> 5.
So it seems that nothing is keeping track of where nulls appear.
Hope someone can take a look, as it is a blocker for my project.
**Environment**: Linux, Java7/Java8
**Reporter**: [Steven Mellinger](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=stevemel)
#### Related issues:
- [filter2 API performance regression](https://github.com/apache/parquet-java/issues/1583) (relates to)
**Note**: *This issue was originally created as [PARQUET-182](https://issues.apache.org/jira/browse/PARQUET-182). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by tracing the nested (x && (y || z)) predicate through FilteredRecordReader, UnboundRecordFilter, and ColumnReader. Inspect RunLenghBitPackingHybridDecoder with an optional column containing null, 0, and 1, and compare decoded values with row positions. Done means filtering returns the correct rows and regression coverage preserves null positions.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100