apache / apache/parquet-java

Schema mismatch for reading Avro from parquet file with old schema version?

Đang mở
#1,628 5 bình luận 0 reaction 0 người được giao Xem trên GitHub
Component: Avro Component: Parquet Priority: Minor Type: bug
Ngôn ngữ chính
Java
Star
3.1k
Fork
1.6k
Merge trung bình
3 ngày 12 giờ
Pull request đã merge (30 ngày)
33

Mô tả

I ran into what looks like a bug in the Parquet Avro reading code, around trying to read a file written with a previous version of a schema with a new, evolved version of the schema.

I'm using Apache Beam's ParquetIO library, which supports passing in schemas to use for "projection" and I was investigating if that would work for me here. However, it didn't work, complaining that my new reader schema had a field that wasn't in the writer schema.

 

I traced this through to a couple places in the parquet-avro code that don't look right to me:

 

First, in `prepareForRead` here:

The `parquetSchema` var comes from `parquetSchema = readContext.getRequestedSchema();` while the `avroSchema` var comes from the parquet file itself with `avroSchema = new Schema.Parser().parse(keyValueMetaData.get(AVRO_SCHEMA_METADATA_KEY));`

I can verify that `parquetSchema` is the schema I'm requesting it be projected to and that `avroSchema` is the schema from the file, but the naming looks backward, shouldn't `parquetSchema` be the one from the parquet file?

Following the stack down, I was hitting this line: https://github.com/apache/parquet-mr/blob/master/parquet-avro/src/main/java/org/apache/parquet/avro/AvroIndexedRecordConverter.java#L91

here it was failing because the `avroSchema` didn't have a field that was in the `parquetSchema`, with the variables assigned in the same way as above. That's the case I was hoping to use this projection for, though - to get the record read with the new reader schema, using the default value from the new schema for the new field. In fact, the comment on line 101 "store defaults for any new Avro fields from avroSchema that are not in the writer schema (parquetSchema)" suggests that the intent was for this to work, but the actual code has the writer schema in avroSchema and the reader schema in parquetSchema.

(Additionally, I'd want this to support schema evolution both for adding an optional field and also removing an old field - so just flipping the names around would result in this still breaking if the reader schema dropped a field from the writer schema...)

Looking to understand if I'm interpreting this correctly, or if there's another path that's intended to be used.

Thank you!

**Environment**: Linux, Apache Beam 2.28.0, Java 11
**Reporter**: [Philip Wilcox](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=philipwilcox)

**Note**: *This issue was originally created as [PARQUET-2055](https://issues.apache.org/jira/browse/PARQUET-2055). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu trong parquet-avro/src/main/java/org/apache/parquet/avro/AvroReadSupport.java tại prepareForRead và lần theo các schema đến AvroIndexedRecordConverter.java quanh dòng 91. So sánh schema chiếu được yêu cầu với schema siêu dữ liệu Avro của tệp, bao gồm cả các chú thích về giá trị mặc định. Được coi là hoàn tất khi hành vi tiến hóa schema dự kiến đối với các trường được thêm và bị xóa được làm rõ hoặc sửa lại, với việc xác thực dựa trên kịch bản Apache Beam đã được báo cáo.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
java
Lĩnh vực
data-engineering
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
35/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.