apache / apache/parquet-java

AvroReadSupport does not support Avro schema resolution

オープン
#1,721 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る
Component: Java Component: Parquet Priority: Major Type: bug
主要言語
Java
スター
3.1k
フォーク
1.6k
平均マージ
3日 12時間
マージ済み PR(30日)
33

説明

Given multiple "different-yet-compatible" Avro-backed Parquet files, a runtime exception will be encountered when trying to merge the metadata values across the files if they are used as input sources for a MapReduce job.

A contrived example of this problem is provided, along with a derived version of `AvroReadSupport` that can correctly handle valid schema resolution/evolution scenarios.

**Illustration of Problem**
A simple Avro schema exists, which contains a single record type that consists of a required String member.

```
{"type":"record","name":"Foo","namespace":"com.tapad.avro","fields":[{"name":"my_field","type":{"type":"string","avro.java.string":"String"}}]}
```

When stored as Parquet-Avro the resulting schema is:

```
message com.tapad.avro.Foo {
required binary my_field (UTF8);
}
```

Data is written to a Parquet-Avro file with the following contents:

```
my_field = aaa
my_field = bbb
```

The schema for the Foo record is then changed so that its String member is made optional, with a default value of null now provided for the String member.

```
{"type":"record","name":"Foo","namespace":"com.tapad.avro","fields":[{"name":"my_field","type":["null",{"type":"string","avro.java.string":"String"}],"default":null}]}

message com.tapad.avro.Foo {
optional binary my_field (UTF8);
}
```

This change adheres to the Avro Schema Resoution rules as found in http://avro.apache.org/docs/current/spec.html#Schema+Resolution.

Data is then written to a new Parquet-Avro file.

```
my_field = ccc
```

When both Parquet-Avro files are used as input to a MapReduce job, wherein the schemas in the data files are considered to be the "writer" schemas and the schema on our job's classpath – in this case, the updated schema – is used as the "reader" schema, the following `RuntimeException` is encountered:

```
Caused by: java.lang.RuntimeException: could not merge metadata: key avro.schema has conflicting values: [{"type":"record","name":"Foo","namespace":"com.tapad.avro","fields":[{"name":"my_field","type":{"type":"string","avro.java.string":"String"}}]}, {"type":"record", "name":"Foo","namespace":"com.tapad.avro","fields":[{"name":"my_field","type":["null",{"type":"string","avro.java.string":"String"}],"default":null}]}]
42 at parquet.hadoop.api.InitContext.getMergedKeyValueMetaData(InitContext.java:67)
43 at parquet.hadoop.api.ReadSupport.init(ReadSupport.java:84)
44 at parquet.hadoop.ParquetInputFormat.getSplits(ParquetInputFormat.java:263)
...
```

**Solution**

Each schema in every data file (the "writer" schemas) should check for schema compatibility with the "reader" schema. If all "writer" schemas are compatible with the "reader" schema, all records in all data files can be migrated to the "reader" schema.

The Apache Avro library provides utilities for performing compatibility checks across schemas and provided is a derived version of `AvroReadSupport` which uses these utilities to successfully process the records in the aforementioned data files when they are both used as input to a MapReduce job.

~~_NOTE: Solution will be provided as a hyperlink to a Github Pull Request_~~ https://github.com/apache/incubator-parquet-mr/pull/107

**Reporter**: [Jeffrey Olchovy](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=jeffo)

**Note**: *This issue was originally created as [PARQUET-171](https://issues.apache.org/jira/browse/PARQUET-171). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

AvroReadSupport と、スタックトレースに示されている ReadSupport メタデータ初期化パスから始め、提供された互換性のある writer schema と reader schema を確認します。両方の Parquet-Avro ファイルを使って MapReduce の入力ケースを再現します。互換性のある writer schema が受け入れられ、メタデータのマージ時に RuntimeException が発生することなく、レコードが reader schema に解決されれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
java
領域
data-engineering
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
30/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。