apache / apache/parquet-java

Add differentiation of nested records with the same name

Ouverte
#2,124 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Component: Avro Component: Parquet Priority: Major Type: bug
Langage dominant
Java
Étoiles
3.1k
Forks
1.6k
Merge moyen
3 j 12 h
PR mergées (30 j)
33

Description

Hello,

While reading back a Parquet file produced with Spark, it appears the schema produced by Parquet-Avro is not valid.

I consider the simple following piece of code:

```Java

ParquetReader reader =

             AvroParquetReader.builder(new org.apache.hadoop.fs.Path(path.toUri())).build();

             System.out.println(reader.read().getSchema());

```

I get a stack lile:

```Java

Exception in thread "main" +org.apache.avro.SchemaParseException+: Can't redefine: value

       at org.apache.avro.Schema$Names.put(+Schema.java:1128+)

       at org.apache.avro.Schema$NamedSchema.writeNameRef(+Schema.java:562+)

       at org.apache.avro.Schema$RecordSchema.toJson(+Schema.java:690+)

       at org.apache.avro.Schema$UnionSchema.toJson(+Schema.java:882+)

       at org.apache.avro.Schema$MapSchema.toJson(+Schema.java:833+)

       at org.apache.avro.Schema$UnionSchema.toJson(+Schema.java:882+)

       at org.apache.avro.Schema$RecordSchema.fieldsToJson(+Schema.java:716+)

       at org.apache.avro.Schema$RecordSchema.toJson(+Schema.java:701+)

       at org.apache.avro.Schema.toString(+Schema.java:324+)

       at org.apache.avro.Schema.toString(+Schema.java:314+)

```

 

The issue seems the same as the one reported in:

 

It have been fixed in Spark-avro within:

In our case, the parquet schema looks like:

```Java

message spark_schema {
optional group calculatedobjectinfomap (MAP) {
repeated group key_value {
required binary key (UTF8);
optional group value {
optional int64 calcobjid;
optional int64 calcobjparentid;
optional binary portfolioname (UTF8);
optional binary portfolioscheme (UTF8);
optional binary calcobjtype (UTF8);
optional binary calcobjmnemonic (UTF8);
optional binary calcobinstrumentype (UTF8);
optional int64 calcobjectqty;
optional binary calcobjboid (UTF8);
optional binary analyticalfoldermnemonic (UTF8);
optional binary calculatedidentifier (UTF8);
optional binary calcobjlevel (UTF8);
optional binary calcobjboidscheme (UTF8);
}
}
}
optional group riskfactorinfomap (MAP) {
repeated group key_value {
required binary key (UTF8);
optional group value {
optional binary riskfactorname (UTF8);
optional binary riskfactortype (UTF8);
optional binary riskfactorrole (UTF8);
}
}
}
}

```
We indeed have 2 Map field with a value fields named 'value'. The name 'value' is defaulted in org.apache.spark.sql.types.MapType.

The fix seems not trivial given current parquet-avro code then I doubt I will be able to craft a valid PR without directions.

Thanks,

**Reporter**: [Benoit Lacelle](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=blasd)

**Note**: *This issue was originally created as [PARQUET-1202](https://issues.apache.org/jira/browse/PARQUET-1202). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Commencez par AvroParquetReader.builder et le chemin de conversion du schéma utilisé lors de la lecture du schéma Parquet montré ; examinez comment MapType fournit le nom de la valeur par défaut. Reproduisez l’échec avec l’exemple de maps imbriquées et comparez l’approche mentionnée dans la pull request 73 de spark-avro. Le travail est terminé lorsque les schémas contenant des noms de records imbriqués répétés se sérialisent sans l’exception "Can't redefine" d’Avro.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
java
Domaine
data-engineering
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
À l'abandon
Clarté
À clarifier
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.