apache / apache/parquet-java

Empty projection returns the wrong number of rows when column index is enabled

Open
#2,702 1 comment 0 reactions 0 assignees View on GitHub
Component: Java Component: Parquet Priority: Major Type: bug
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

Discovered in Spark, when returning an empty projection from a Parquet file with filter pushdown enabled (typically when doing filter + count), Parquet-Mr returns a wrong number of rows with column index enabled. When the column index feature is disabled, the result is correct.

 

This happens due to the following:
1. ParquetFileReader::getFilteredRowCount() ( selects row ranges to calculate the row count when column index is enabled.
1. In ColumnIndexFilter ( we filter row ranges and pass the set of paths which in this case is empty.
1. When evaluating the filter, if the column path is not in the set, we would return an empty list of rows ([https://github.com/apache/parquet-mr/blob/0819356a9dafd2ca07c5eab68e2bffeddc3bd3d9/parquet-column/src/main/java/org/apache/parquet/internal/filter2/columnindex/ColumnIndexFilter.java#L178)](https://github.com/apache/parquet-mr/blob/0819356a9dafd2ca07c5eab68e2bffeddc3bd3d9/parquet-column/src/main/java/org/apache/parquet/internal/filter2/columnindex/ColumnIndexFilter.java#L178).) which is always the case for an empty projection.
1. This results in the incorrect number of records reported by the library.

I will provide the full repro later.

 

 

**Reporter**: [Ivan Sadikov](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=ivan.sadikov)
#### Related issues:
- [Filtered parquet data frame count() and show() produce inconsistent results when spark.sql.parquet.filterPushdown is true](https://issues.apache.org/jira/browse/SPARK-39833) (is related to)

**Note**: *This issue was originally created as [PARQUET-2170](https://issues.apache.org/jira/browse/PARQUET-2170). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with ParquetFileReader::getFilteredRowCount() and ColumnIndexFilter, following the empty column-path set through row-range evaluation. Reproduce the Spark filter-plus-count case once the full repro is available, then verify that an empty projection with column indexes and filter pushdown reports the correct row count.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.