apache / apache/parquet-java

`RequiresFallback.isCompressionSatisfying` is too aggressive with current default page size

オープン
#3,479 コメント 1 件 リアクション 1 件 担当者 0 名 GitHub で見る
Type: enhancement
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
3日 12時間
マージ済み PR(30日)
33

説明

### Describe the enhancement requested

An issue recently was brought up in arrow-rs (https://github.com/apache/arrow-rs/pull/9700) which brought to my attention the existence of `isCompressionSatisfying` in the `RequiresFallback` interface. In short, after accumulating a page worth of data, `isCompressionSatisfying` is called to see if dictionary encoding is actually compressing the data at all, and if not, then the encoder falls back immediately to the fallback encoder. As far as I could determine, this behavior was introduced very early on, before the advent of the page indexes, so IIRC the page size would have been significantly larger. With page indexes, however, this function is now called after only 20000 rows have been processed. A column with a moderate cardinality might not yet have produced enough repeating values to lead this function to conclude it's best to continue using a dictionary.

For example, a dataframe with an int64 column consisting of one million values mod'd with 32768 will end up ditching dictionary encoding completely, and produce a column chunk of 8.4MB. If the page row count is bumped up to 128k, then dictionary encoding is used throughout and the resultant column chunk is only 2.2MB.

Sadly, it does not appear that this behavior is configurable, so short of increasing the page row count, its behavior cannot be modified.

I can see the need for this type of heuristic, but I think it needs to be modified in light of the current defaults resulting in far too few samples with which to determine if dictionary encoding is beneficial or not. If collecting more samples before falling back is not practical, there should at least be a configuration setting to disable this check.

### Component(s)

Core

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

RequiresFallback.isCompressionSatisfying エントリポイントから開始し、ページインデックス、20,000 行のデフォルト値、辞書フォールバックがどのように相互作用するかを追跡します。報告された 20,000 行と 128,000 行の結果を比較し、合意された変更がサンプリングヒューリスティックの改訂なのか、設定オプションなのかを判断します。中程度のカーディナリティのデータに対する動作がカバーされれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java
領域
data-engineering
issue の種類
機能追加
難易度
4/5
見積もり時間
3〜5日
活発さ
静か
明瞭さ
おおむね明確
初心者へのやさしさ
45/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。