Invalid SchemaException for UUID while using AvroParquetWriter
- Ngôn ngữ chính
- Java
- Star
- 3.1k
- Fork
- 1.6k
- Merge trung bình
- 3 ngày 12 giờ
- Pull request đã merge (30 ngày)
- 33
Mô tả
Hi,
I am getting org.apache.parquet.schema.InvalidSchemaException: Cannot write a schema with an empty group: optional group id {} while I include a UUID field on my POJO object. Without UUID everything worked fine. I have seen Parquet suports UUID as part of [#PR-71] on 2.4 release.
But I am getting InvalidSchemaException on UUID. Is there anything that I am missing or its a known issue?
**My setup details:**
**gradle dependency :**
dependencies
{ compile group: 'org.springframework.boot', name: 'spring-boot-starter' compile group: 'org.projectlombok', name: 'lombok', version: '1.16.6' compile group: 'com.amazonaws', name: 'aws-java-sdk-bundle', version: '1.11.271' compile group: 'org.apache.parquet', name: 'parquet-avro', version: '1.10.1' compile group: 'org.apache.hadoop', name: 'hadoop-common', version: '3.1.1' compile group: 'org.apache.hadoop', name: 'hadoop-aws', version: '3.1.1' compile group: 'org.apache.hadoop', name: 'hadoop-client', version: '3.1.1' compile group: 'joda-time', name: 'joda-time' compile group: 'com.fasterxml.jackson.core', name: 'jackson-databind', version: '2.6.5' compile group: 'com.fasterxml.jackson.datatype', name: 'jackson-datatype-joda', version: '2.6.5' }
**Model used:**
@Data
public class Employee
{ private UUID id; private String name; private int age; private Address address; }
@Data
public class Address
{ private String streetName; private String city; private Zip zip; }
@Data
public class Zip
{ private int zip; private int ext; }
**My Serializer Code:**
public void serialize(List inputDataToSerialize, CompressionCodecName compressionCodecName) throws IOException {
Path path = new Path("s3a://parquetpoc/data_"+compressionCodecName+".parquet");
Class clazz = inputDataToSerialize.get(0).getClass();
try (ParquetWriter writer = AvroParquetWriter.builder(path)
.withSchema(ReflectData.AllowNull.get().getSchema(clazz)) // generate nullable fields
.withDataModel(ReflectData.get())
.withConf(parquetConfiguration)
.withCompressionCodec(compressionCodecName)
.withWriteMode(OVERWRITE)
.withWriterVersion(ParquetProperties.WriterVersion.PARQUET_2_0)
.build()) {
for (D input : inputDataToSerialize)
{ writer.write(input); }
}
}
private List **getInputDataToSerialize**(){
Address address = new Address();
address.setStreetName("Murry Ridge Dr");
address.setCity("Murrysville");
Zip zip = new Zip();
zip.setZip(15668);
zip.setExt(1234);
address.setZip(zip);
List employees = new ArrayList<>();
IntStream.range(0, 100000).forEach(i->
{ Employee employee = new Employee(); // employee.setId(UUID.randomUUID()); employee.setAge(20); employee.setName("Test"+i); employee.setAddress(address); employees.add(employee); }
);
return employees;
}
_\*\*Where generic Type D is Employee_
**Reporter**: [Felix Kizhakkel Jose](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=FelixKJose) / @FelixKJose
**Note**: *This issue was originally created as [PARQUET-1679](https://issues.apache.org/jira/browse/PARQUET-1679). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Hướng nghiên cứu
Tái hiện lỗi với các model Employee, Address và Zip, bao gồm cả trường UUID, cùng lời gọi AvroParquetWriter.builder sử dụng ReflectData.AllowNull.get().getSchema(clazz). So sánh schema được tạo ra với phiên bản hoạt động đúng không có UUID. Được coi là hoàn tất khi xác định được vấn đề tạo schema của UUID và xác minh hành vi đã sửa bằng một regression test hoặc kết quả tương thích được ghi lại.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- java
- Lĩnh vực
- data-engineering
- Loại issue
- Lỗi
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Cần làm rõ
- Mức phù hợp với người mới
- 35/100