apache / apache/parquet-java

Avro's isElementType() change breaks the reading of some parquet(1.8.1) files

Open
#2,381 13 comments 0 reactions 0 assignees View on GitHub
Component: Avro Component: Parquet Priority: Critical Type: enhancement
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

When using the Avro schema below to write a parquet(1.8.1) file and then read back by using parquet 1.10.1 without passing any schema, the reading throws an exception "XXX is not a group" . Reading through parquet 1.8.1 is fine. 

           {

              "name": "phones",

              "type": [

                "null",

                {

                  "type": "array",

                  "items": {

                    "type": "record",

                    "name": "phones_items",

                    "fields": [

                      

{                         "name": "phone_number",                         "type": [                           "null",                           "string"                         ],                         "default": null                       }

                    ]

                  }

                }

              ],

              "default": null

            }

The code to read is as below 

     val reader = AvroParquetReader._builder_[SomeRecordType](parquetPath).withConf(**new**   Configuration).build()

    reader.read()

PARQUET-651 changed the method isElementType() by relying on Avro's checkReaderWriterCompatibility() to check the compatibility. However, checkReaderWriterCompatibility() consider the ParquetSchema and the AvroSchema(converted from File schema) as not compatible(the name in avro schema is ‘phones_items’, but the name is ‘array’ in Parquet schema, hence not compatible) . Hence return false and caused the “phone_number” field in the above schema to be considered as group type which is not true. Then the exception throws as .asGroupType(). 

I didn’t try writing via parquet 1.10.1 would reproduce the same problem or not. But it could because the translation of Avro schema to Parquet schema is not changed(didn’t verify yet). 

 I hesitate to revert PARQUET-651 because it solved several problems. I would like to hear the community's thoughts on it. 

**Reporter**: [Xinli Shang](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=shangx@uber.com) / @shangxinli
**Assignee**: [Xinli Shang](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=shangx@uber.com) / @shangxinli
#### Related issues:
- [Parquet-avro fails to decode array of record with a single field name "element" correctly](https://github.com/apache/parquet-java/issues/1976) (causes)

**Note**: *This issue was originally created as [PARQUET-1681](https://issues.apache.org/jira/browse/PARQUET-1681). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Reproduce the failure with the supplied Avro schema, parquet 1.8.1 output, parquet 1.10.1, and AvroParquetReader, then inspect the isElementType() logic changed by PARQUET-651 and its compatibility check. Done means older files with this array-of-record schema can be read without the “not a group” exception, with the behavior covered by a regression test if the repository provides one.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.