apache / apache/parquet-java

Parquet Java Serialization is very slow

Đang mở
#1,569 2 bình luận 0 reaction 0 người được giao Xem trên GitHub
Component: Avro Component: Java Component: Parquet Priority: Major Type: bug
Ngôn ngữ chính
Java
Star
3.1k
Fork
1.6k
Merge trung bình
3 ngày 12 giờ
Pull request đã merge (30 ngày)
33

Mô tả

Hi,
I am doing a POC to compare different data formats and its performance in terms of serialization/deserialization speed, storage size, compatibility between different language etc. 
When I try to serialize a simple java object to parquet file,  it takes _\*6-7 seconds\*_ vs same object's serialization to JSON is **_100 milliseconds._**

Could you help me to resolve this issue?

+**My Configuration and code snippet:**
**Gradle dependencies**
dependencies

{ compile group: 'org.springframework.boot', name: 'spring-boot-starter' compile group: 'org.projectlombok', name: 'lombok', version: '1.16.6' compile group: 'com.amazonaws', name: 'aws-java-sdk-bundle', version: '1.11.271' compile group: 'org.apache.parquet', name: 'parquet-avro', version: '1.10.0' compile group: 'org.apache.hadoop', name: 'hadoop-common', version: '3.1.1' compile group: 'org.apache.hadoop', name: 'hadoop-aws', version: '3.1.1' compile group: 'org.apache.hadoop', name: 'hadoop-client', version: '3.1.1' compile group: 'joda-time', name: 'joda-time' compile group: 'com.fasterxml.jackson.core', name: 'jackson-databind', version: '2.6.5' compile group: 'com.fasterxml.jackson.datatype', name: 'jackson-datatype-joda', version: '2.6.5' }

**Code snippet:**+

public void serialize(List inputDataToSerialize, CompressionCodecName compressionCodecName) throws IOException {

Path path = new Path("s3a://parquetpoc/data_"+compressionCodecName+".parquet");
Path path1 = new Path("/Downloads/data_"+compressionCodecName+".parquet");
Class clazz = inputDataToSerialize.get(0).getClass();

try (ParquetWriter writer = **AvroParquetWriter.**builder(path1)
.withSchema(ReflectData.AllowNull.get().getSchema(clazz)) // generate nullable fields
.withDataModel(ReflectData.get())
.withConf(parquetConfiguration)
.withCompressionCodec(compressionCodecName)
.withWriteMode(OVERWRITE)
.withWriterVersion(ParquetProperties.WriterVersion.PARQUET_2_0)
.build()) {

for (D input : inputDataToSerialize)

{ writer.write(input); }

}
}

+**Model Used:**
@Data
public class Employee

{ //private UUID id; private String name; private int age; private Address address; }

@Data
public class Address

{ private String streetName; private String city; private Zip zip; }

@Data
public class Zip

{ private int zip; private int ext; }

 

private List **getInputDataToSerialize**(){
Address address = new Address();
address.setStreetName("Murry Ridge Dr");
address.setCity("Murrysville");
Zip zip = new Zip();
zip.setZip(15668);
zip.setExt(1234);

address.setZip(zip);

List employees = new ArrayList<>();

IntStream.range(0, 100000).forEach(i->{
Employee employee = new Employee();
// employee.setId(UUID.randomUUID());
employee.setAge(20);
employee.setName("Test"+i);
employee.setAddress(address);
employees.add(employee);
});
return employees;
}

**Note:**
**I have tried to save the data into local file system as well as AWS S3, but both are having same result - very slow.**

**Reporter**: [Felix Kizhakkel Jose](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=FelixKJose) / @FelixKJose

**Note**: *This issue was originally created as [PARQUET-1680](https://issues.apache.org/jira/browse/PARQUET-1680). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu bằng việc tái hiện benchmark xoay quanh AvroParquetWriter.builder, ReflectData.AllowNull.get().getSchema(clazz) và writer.write bằng model Employee được cung cấp cùng input gồm 100,000 bản ghi. So sánh việc tạo schema, khởi tạo writer và ghi trên đường dẫn cục bộ và đường dẫn S3; được xem là hoàn tất khi đã xác định được nguồn gây ra độ trễ lớn nhất và chứng minh được một cải thiện qua đo lường.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
aws, hadoop, java
Lĩnh vực
data-engineering, performance
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
25/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.