apache / apache/datafusion

Initialize TopK from file / rowgroup / .. statistics

Open
#21,691 0 comments 0 reactions 1 assignee Claimed by @zhuqi-lucas View on GitHub
performance
Dominant language
Rust
Stars
9.3k
Forks
2.4k
Avg merge
3d 7h
Merged PRs (30d)
344

Description

We could initialize the TopK statistics from column stats (at least for single columns) and make the initial threshold much tighter based on min/max statistics (at file / rowgroup/page level):

* We have a file/rowgroup with more than K (from TopK) amount of rows
* We have a single sort column (directly after scan)
* We can initialize/update the TopK using max (or min) statistics
* Also, if the new bound is smaller / bigger than the current TopK, we could update it to the tighter bound

This I think might help making initial threshold much tighter instead of having to read all the first row groups using not-initialized TopK.

_Originally posted by @Dandandan in https://github.com/apache/datafusion/issues/21580#issuecomment-4266553697_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.