apache / apache/parquet-java

Cannot read row group larger than 2GB

未关闭
#2,057 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: Java Component: Parquet Priority: Major Type: bug
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

Parquet MR 1.8.2 does not support reading row groups which are larger than 2 GB. See:https://github.com/apache/parquet-mr/blob/parquet-1.8.x/parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java#L1064

We are seeing this when writing skewed records. This throws off the estimation of the memory check interval in the InternalParquetRecordWriter. The following spark code illustrates this:
```
/**
* Create a data frame that will make parquet write a file with a row group larger than 2 GB. Parquet
* only checks the size of the row group after writing a number of records. This number is based on
* average row size of the already written records. This is problematic in the following scenario:
* - The initial (100) records in the record group are relatively small.
* - The InternalParquetRecordWriter checks if it needs to write to disk (it should not), it assumes
* that the remaining records have a similar size, and (greatly) increases the check interval (usually
* to 10000).
* - The remaining records are much larger then expected, making the row group larger than 2 GB (which
* makes reading the row group impossible).
*
* The data frame below illustrates such a scenario. This creates a row group of approximately 4GB.
*/
val badDf = spark.range(0, 2200, 1, 1).mapPartitions { iterator =>
var i = 0
val random = new scala.util.Random(42)
val buffer = new Array[Char](750000)
iterator.map { id =>
// the first 200 records have a length of 1K and the remaining 2000 have a length of 750K.
val numChars = if (i < 200) 1000 else 750000
i += 1

// create a random array
var j = 0
while (j < numChars) {
// Generate a char (borrowed from scala.util.Random)
buffer(j) = (random.nextInt(0xD800 - 1) + 1).toChar
j += 1
}

// create a string: the string constructor will copy the buffer.
new String(buffer, 0, numChars)
}
}
badDf.write.parquet("somefile")
val corruptedDf = spark.read.parquet("somefile")
corruptedDf.select(count(lit(1)), max(length($"value"))).show()
```
The latter fails with the following exception:
```
java.lang.NegativeArraySizeException
at org.apache.parquet.hadoop.ParquetFileReader$ConsecutiveChunkList.readAll(ParquetFileReader.java:1064)
at org.apache.parquet.hadoop.ParquetFileReader.readNextRowGroup(ParquetFileReader.java:698)
...
```

-This seems to be fixed by commit https://github.com/apache/parquet-mr/commit/6b605a4ea05b66e1a6bf843353abcb4834a4ced8 in parquet 1.9.x. Is there any chance that we can fix this in 1.8.x?-

**Reporter**: [Herman van Hövell](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=hvanhovell)

**Note**: *This issue was originally created as [PARQUET-980](https://issues.apache.org/jira/browse/PARQUET-980). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 parquet-hadoop/src/main/java/org/apache/parquet/hadoop/ParquetFileReader.java 第 1064 行开始,并将其行为与 commit 6b605a4ea05b66e1a6bf843353abcb4834a4ced8 进行比较。使用报告中的 Spark 示例重现该问题,然后验证 parquet 1.8.x 能够读取生成的 row group,且不会出现 NegativeArraySizeException。

由索引模型根据 Issue 内容生成。

评估

技术栈
java, scala
领域
data-engineering
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
描述清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。