apache / apache/parquet-java

record count for row group size check configurable

オープン
#2,740 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
Component: Java Component: Parquet Priority: Major Type: enhancement
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
3日 12時間
マージ済み PR(30日)
33

説明

 org.apache.parquet.hadoop.InternalParquetRecordWriter#checkBlockSizeReached
```java

 private void checkBlockSizeReached() throws IOException {
    if (recordCount >= recordCountForNextMemCheck) { // checking the memory size is relatively expensive, so let's not do it for every record.
      long memSize = columnStore.getBufferedSize();
      long recordSize = memSize / recordCount;
      // flush the row group if it is within ~2 records of the limit
      // it is much better to be slightly under size than to be over at all
      if (memSize > (nextRowGroupSize - 2 * recordSize)) {
        LOG.info("mem size {} > {}: flushing {} records to disk.", memSize, nextRowGroupSize, recordCount);
        flushRowGroupToStore();
        initStore();
        recordCountForNextMemCheck = min(max(MINIMUM_RECORD_COUNT_FOR_CHECK, recordCount / 2), MAXIMUM_RECORD_COUNT_FOR_CHECK);
        this.lastRowGroupEndPos = parquetFileWriter.getPos();
      } else {
        recordCountForNextMemCheck = min(
            max(MINIMUM_RECORD_COUNT_FOR_CHECK, (recordCount + (long)(nextRowGroupSize / ((float)recordSize))) / 2), // will check halfway
            recordCount + MAXIMUM_RECORD_COUNT_FOR_CHECK // will not look more than max records ahead
            );
        LOG.debug("Checked mem at {} will check again at: {}", recordCount, recordCountForNextMemCheck);
      }
    }
  }
```
in this code,if the block size is small ,for example 8M,and the first 100 lines record size is small and  after 100 lines the record size is big,it will cause big row group,in our real scene,it will more than 64M. So i think the size for block check can configurable.

**Reporter**: [xjlem](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=xjlem)

**Note**: *This issue was originally created as [PARQUET-2242](https://issues.apache.org/jira/browse/PARQUET-2242). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

org.apache.parquet.hadoop.InternalParquetRecordWriter#checkBlockSizeReached から始め、nextRowGroupSize と recordCountForNextMemCheck がどのように初期化され、使用されているかを追跡します。既存の構成経路と関連する writer テストを特定し、メモリチェックのサイズ設定を構成可能にし、報告されている oversized row-group の動作を防止できることを完了条件として定義します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java
領域
data-engineering
issue の種類
機能追加
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
42/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。