Parquet File not readable by Google big query (works with Spark)
- 主要言語
- Java
- スター
- 3.1k
- フォーク
- 1.6k
- 平均マージ
- 3日 12時間
- マージ済み PR(30日)
- 33
説明
Hi
I'm trying to write Avro message to parquet on GCS. These parquet should be query by big query engine who support now parquet.
To do this I'm using Secor a kafka log persister tools from pinterest.
First I didn't notice any problem using Spark the same file can be read without any problem all is working perfect.
Now using Big query bring and error like this :
Error while reading table: , error message: Read less values than expected: Actual: 29333, Expected: 33827. Row group: 0, Column: , File:
After investigation using parquet-tools I figured out that in parquet there is metadata regarding number total of unique values for each columns eg from parquet-tools
page 0: DLE:BIT_PACKED RLE:BIT_PACKED [more]... CRC:[PAGE CORRUPT] VC:547
So the VC value indicate that the total number of unique value in the file is 547.
Now when make a spark SQL like SELECT DISTINCT COUNT(column) FROM ... I get 421 mean this number in the metadata is incorrect.
So what is not a problem for Spark to read is a blocking problem for Big data because it relies on these values and found it incorrect.
Is there any configuration of the writer that can prevent these errors in the metadata ? Or is it a normal behavior that should be a problem ?
Thanks
**Environment**: [secor|https://github.com/pinterest/secor]
GCP
Big Query google cloud
Parquet writer 1.11
**Reporter**: [Richard Grossman](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=richiesgr)
**Note**: *This issue was originally created as [PARQUET-1946](https://issues.apache.org/jira/browse/PARQUET-1946). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
調査の方向性
リポジトリのファイルもテストも指定されていません。まず Parquet writer 1.11 で Secor の出力を再現し、BigQuery で読み取ってから、報告されたメタデータを Spark および parquet-tools の結果と比較してください。非互換性を特定し、対象を絞った回帰ケースまたは文書化された制限事項によって確認できれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- gcp, google-cloud, java
- 領域
- data-engineering, databases
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 30/100