apache / apache/arrow-rs

Optimized decoding of Parquet Statistics, `null_pages` and `null_counts`

Open
#9,296 10 comments 0 reactions 0 assignees View on GitHub
enhancement performance
Dominant language
Rust
Stars
3.6k
Forks
1.3k
Avg merge
2d 18h
Merged PRs (30d)
169

Description

**Is your feature request related to a problem or challenge? Please describe what you are trying to do.**
Currently, if using statistics, a lot of time can be spent decoding/summarizing the statistics from the `ValueStatistics` / `Statistics` structs (which are large / inefficient structs).
In DataFusion this can sometimes take as much time running the query (or more if the query can be answered from statistics directly).

**Describe the solution you'd like**
We should consider decoding the statistics into a columnar format (values + null bitmap (if needed)) directly, avoiding needing to convert this later (and possibly help decoding as well a bit as well as memory usage).

Looking at `ColumnIndex`:

```
pub struct ColumnIndex {
pub(crate) null_pages: Vec,
pub(crate) boundary_order: BoundaryOrder,
pub(crate) null_counts: Option>,
pub(crate) repetition_level_histograms: Option>,
pub(crate) definition_level_histograms: Option>,
}
```

* `null_pages`: this currently is a `Vec` (true is null, false is non-null), it would be better to save this as a `NullBuffer` or similar, where `true` means valid and `false` means invalid - this would make it possible to copy the null bitmap without conversion
* `null_counts`: Option>: it would be better to have this as a `Int64Array` or similar (Or preferably even `Uint64Array` if we can do the conversion earlier)

**Describe alternatives you've considered**

**Additional context**

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.