apache / apache/parquet-java

Support lazy materialization of row groups in ParquetFileReader

Abierto
#2,884 3 comentarios 0 reacciones 0 asignados Ver en GitHub
Component: Parquet Priority: Major Type: enhancement
Lenguaje dominante
Java
Estrellas
3.1k
Forks
1.6k
Merge medio
3 d 12 h
PR fusionados (30 d)
33

Descripción

Motivation: The current behavior of ParquetFilterReader#readNextRowGroup is to eagerly enumerate all chunks in the row group, then read all pages in the chunk. For distributed data workloads, this can cause significant memory pressure, particularly for use cases that require the colocation of multiple Parquet files on a single worker.

 

Proposal: A Parquet Configuration option that enables lazy row group reading, i.e., only a page at a time (plus whatever header is necessary to read that header). The Configuration option could be either a flag, or an int value for how many pages/page bytes to buffer at a time.

 

I think this could be accomplished by modifying [ParquetFileReader#readAllPages](https://github.com/apache/parquet-mr/blob/cf294a35ca4b635f44357fe4c170f3b66870dc6e/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java#L1727) to re-implement pagesInChunk as an Iterator, rather than a List. Then, ColumnChunkPageReader could parse the Configuration option above and decide whether to fully materialize the iterator or not.

 

I'm happy to try to create a draft/branch for this to get some early feedback on the idea!

**Reporter**: [Claire McGinty](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=clairemcginty) / @clairemcginty
#### PRs and other links:
- [GitHub Pull Request #1293](https://github.com/apache/parquet-mr/pull/1293)

**Note**: *This issue was originally created as [PARQUET-2443](https://issues.apache.org/jira/browse/PARQUET-2443). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Comienza con ParquetFileReader#readAllPages en parquet-hadoop y con el comportamiento de ColumnChunkPageReader descrito en la propuesta; revisa Pull Request #1293 para consultar el trabajo existente. Se considera terminado cuando una opción de configuración de Parquet admite la materialización perezosa de las páginas de los row groups, al tiempo que conserva el comportamiento eager existente cuando está deshabilitada.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
java
Área
data-engineering
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.