AbsaOSS / AbsaOSS/spark-partition-sizing
In `DataFramePartitioner.DataFrameFunctions.repartitionByRecordCount` use max instead of average to decide about partitoing
未关闭
enhancement
- 主要语言
- Scala
- 星标
- 9
- 派生
- 2
- PR 合并指标
- 30 天内没有已合并 PR
描述
## Background
In `DataFramePartitioner.DataFrameFunctions.repartitionByRecordCount` the limit is compared against the average of records per partition. That can be under the limit while some partitions are still over the limit.
## Feature
Change the method to decide about portioning based on the `max` of record counts per partition if to do or not the repartitioing.
Or make this parametrized or a dedicated method.
贡献指南
这个仓库没有索引到贡献指南
评估
这个 Issue 还没有评估数据。