Out of Memory when reading large parquet file
- Langage dominant
- Java
- Étoiles
- 3.1k
- Forks
- 1.6k
- Merge moyen
- 3 j 12 h
- PR mergées (30 j)
- 33
Description
Hi,
We are successfully reading parquet files block by block, and are running into a JVM out of memory issue in a certain edge case. Consider the following scenario:
Parquet file has one column and one block and is 10 GB
Our JVM is 5 GB
Is there any way to read such a file? Below is our implementation/stack trace
```java
Caused by: java.lang.OutOfMemoryError: Java heap space
at org.apache.parquet.hadoop.ParquetFileReader$ConsecutiveChunkList.readAll(ParquetFileReader.java:778)
at org.apache.parquet.hadoop.ParquetFileReader.readNextRowGroup(ParquetFileReader.java:511)
try {
ParquetMetadata readFooter = ParquetFileReader.readFooter(hfsConfig, path,
ParquetMetadataConverter.NO_FILTER);
MessageType schema = readFooter.getFileMetaData().getSchema();
long a = readFooter.getBlocks().stream().
reduce(0L, (left, right) -> left >
right.getTotalByteSize() ? left : right.getTotalByteSize(),
(leftl, rightl) -> leftl > rightl ? leftl : rightl);
for (BlockMetaData block : readFooter.getBlocks()) {
try {
fileReader = new ParquetFileReader(hfsConfig,
readFooter.getFileMetaData(), path, Collections
.singletonList(block), schema.getColumns());
PageReadStore pages;
while (null != (pages = fileReader.readNextRowGroup())) {
//exception gets thrown here on blocks larger than jvm memory
final long rows = pages.getRowCount();
final MessageColumnIO columnIO = new
ColumnIOFactory().getColumnIO(schema);
final RecordReader recordReader =
columnIO.getRecordReader(pages, new GroupRecordConverter(schema));
for (int i = 0; i < rows; i++) {
final Group group = recordReader.read();
int fieldCount = group.getType().getFieldCount();
for (int field = 0; field < fieldCount; field++) {
int valueCount = group.getFieldRepetitionCount(field);
Type fieldType = group.getType().getType(field);
String fieldName = fieldType.getName();
for (int index = 0; index < valueCount; index++) {
// Process data
}
}
}
}
} catch (IOException e) {
...
} finally {
...
}
}
```
**Reporter**: [Ryan Sachs](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=sachsry)
**Note**: *This issue was originally created as [PARQUET-1359](https://issues.apache.org/jira/browse/PARQUET-1359). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Piste de recherche
Commencez par reproduire la lecture Parquet d’une seule colonne de 10 Go avec une JVM de 5 Go, puis suivez ParquetFileReader.readNextRowGroup jusqu’à ConsecutiveChunkList.readAll, les emplacements indiqués dans la stack trace. Déterminez s’il est attendu que le reader gère un row group plus grand que le heap, et définissez l’achèvement comme un comportement documenté ou testé qui évite l’échec Out-of-Memory signalé.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- java
- Domaine
- data-engineering
- Type d'issue
- Bug
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- À l'abandon
- Clarté
- À clarifier
- Accessibilité débutants
- 25/100