`RequiresFallback.isCompressionSatisfying` is too aggressive with current default page size
- 主要语言
- Java
- 星标
- 3.1k
- 派生
- 1.6k
- 平均合并
- 3 天 12 小时
- 30 天内合并 PR
- 33
描述
### Describe the enhancement requested
An issue recently was brought up in arrow-rs (https://github.com/apache/arrow-rs/pull/9700) which brought to my attention the existence of `isCompressionSatisfying` in the `RequiresFallback` interface. In short, after accumulating a page worth of data, `isCompressionSatisfying` is called to see if dictionary encoding is actually compressing the data at all, and if not, then the encoder falls back immediately to the fallback encoder. As far as I could determine, this behavior was introduced very early on, before the advent of the page indexes, so IIRC the page size would have been significantly larger. With page indexes, however, this function is now called after only 20000 rows have been processed. A column with a moderate cardinality might not yet have produced enough repeating values to lead this function to conclude it's best to continue using a dictionary.
For example, a dataframe with an int64 column consisting of one million values mod'd with 32768 will end up ditching dictionary encoding completely, and produce a column chunk of 8.4MB. If the page row count is bumped up to 128k, then dictionary encoding is used throughout and the resultant column chunk is only 2.2MB.
Sadly, it does not appear that this behavior is configurable, so short of increasing the page row count, its behavior cannot be modified.
I can see the need for this type of heuristic, but I think it needs to be modified in light of the current defaults resulting in far too few samples with which to determine if dictionary encoding is beneficial or not. If collecting more samples before falling back is not practical, there should at least be a configuration setting to disable this check.
### Component(s)
Core
贡献指南
这个仓库没有索引到贡献指南
调研方向
从 RequiresFallback.isCompressionSatisfying 入口开始,跟踪页面索引、20,000 行默认值和字典回退之间的交互。比较报告中的 20,000 行和 128,000 行结果,然后确定约定的更改是修订采样启发式算法,还是配置选项;对于中等基数的数据覆盖该行为即表示完成。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- java
- 领域
- data-engineering
- Issue 类型
- 功能
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 冷清
- 描述清晰度
- 基本清楚
- 新手友好度
- 45/100