apache / apache/parquet-java

ProtoReader does not iterate over the parquet file correctly.

オープン
#2,497 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
Component: Java Component: Parquet Priority: Major Type: bug
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
3日 12時間
マージ済み PR(30日)
33

説明

The `ProtoParquetReader` does not iterate over the parquet file correctly, but it gets stuck in the first element and keeps reading as many times as elements the file contained.

In my Scala example I am just reading from a local file that I know for sure it contains right data.

```

val hadoopCOnf = new Configuration()

val outfile: String = genTemporaryFile()

val r: ParquetReader[Event.Builder] = {
ProtoParquetReader.builder[Event.Builder](new Path(outfile)).withConf(hadoopCOnf).build()
}

```

Notice that the proto schema that I am using is generated from 

The generated proto implements com.google.protobuf.GeneratedMessageV3.

See an example on how the ProtoParquetReader is created line(65 and 69): 

and here how it is used (notice that is defined for only one record, the same one for multiple records would fail) 

 

**Reporter**: [Pau Alarcon ](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=paualarco)

**Note**: *This issue was originally created as [PARQUET-1871](https://issues.apache.org/jira/browse/PARQUET-1871). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

ProtoParquetReader から始め、monix-connect の ProtoParquetFixture.scala の 65~69 行付近での構築方法と、ProtoParquetSpec.scala の 83 行付近での使用方法を比較します。複数のレコードで動作を再現し、反復処理によって最初のレコードが繰り返されるのではなく、後続のレコードが返されることを確認します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java, scala
領域
data-engineering
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。