apache / apache/parquet-java

Document recommendations for block size and page size given an expected number of writers

Abierto
#1,709 3 comentarios 0 reacciones 0 asignados Ver en GitHub
Component: Parquet Priority: Major Type: enhancement
Lenguaje dominante
Java
Estrellas
3.1k
Forks
1.6k
Merge medio
3 d 12 h
PR fusionados (30 d)
33

Descripción

I sent a mail on dev list but I seem to have a problem with email on dev list so opening a bug here .

I am on a multithreaded system where there are M threads , each thread creating an independent parquet writer and writing on the hdfs in its own independent files . I have a finite amount of RAM say R .

Now when I created parquet writer using default block and page size i get heap error (no memory ) on my set up . so I reduced my block size and page size to very low and my system stopped giving me these out of memory errors and started writing the file correctly . I am able to read these files correctly as well .

I should not have to make the memory low and parquet should automatically make sure i do not get these errors .

But in case i have to keep track of the memory my question is as follows.

Now keeping these values very less is not a recommended practice as i would loose on performance . I am particularly concerned about write performance . What math formula do you recommend that I should use to find correct blockSize , pageSize to be passed to the parquet constructor to have the right WRITE performance . ie how can i decide what should be the right blockSize , pageSize for a parquet writer given that i have M threads and total RAM memory available is R . I don't understand dictionaryPageSize need and in case i need to bother about that as well kindly let me know but i have kept enableDictionary flag as false .

I am using the bellow constructor .
public More ...ParquetWriter(
162 Path file,
163 WriteSupport writeSupport,
164 CompressionCodecName compressionCodecName,
165 int blockSize,
166 int pageSize,
167 int dictionaryPageSize,
168 boolean enableDictionary,
169 boolean validating,
170 WriterVersion writerVersion,
171 Configuration conf) throws IOException {

**Reporter**: [Manish Agarwal](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=manish.agarwal)

**Note**: *This issue was originally created as [PARQUET-156](https://issues.apache.org/jira/browse/PARQUET-156). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Comience con el constructor de ParquetWriter y sus parámetros blockSize, pageSize, dictionaryPageSize y enableDictionary. Determine qué recomendaciones puede respaldar el proyecto para M writers y una memoria total R, y documente después las compensaciones y las consideraciones sobre las páginas de diccionario. Se considera terminado cuando los usuarios dispongan de recomendaciones prácticas para dimensionar la configuración, centradas en el rendimiento de escritura.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
java
Área
data-engineering
Tipo de issue
Documentación
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.