Unnecessary getFileStatus() calls on all part-files in ParquetInputFormat.getSplits
- 主要语言
- Java
- 星标
- 3.1k
- 派生
- 1.6k
- 平均合并
- 3 天 12 小时
- 30 天内合并 PR
- 33
描述
When testing Spark SQL Parquet support, we found that accessing large Parquet files located in S3 can be very slow. To be more specific, we have a S3 Parquet file with over 3,000 part-files, calling `ParquetInputFormat.getSplits` on it takes several minutes. (We were accessing this file from our office network rather than AWS.)
After some investigation, we found that `ParquetInputFormat.getSplits` is trying to call `getFileStatus()` on all part-files one by one sequentially ([here](https://github.com/apache/incubator-parquet-mr/blob/parquet-1.5.0/parquet-hadoop/src/main/java/parquet/hadoop/ParquetInputFormat.java#L370)). And in the case of S3, each `getFileStatus()` call issues an HTTP request and wait for the reply in a blocking manner, which is considerably expensive.
Actually all these `FileStatus` objects have already been fetched when footers are retrieved ([here](https://github.com/apache/incubator-parquet-mr/blob/parquet-1.5.0/parquet-hadoop/src/main/java/parquet/hadoop/ParquetInputFormat.java#L443)). Caching these `FileStatus` objects can greatly improve our S3 case (reduced from over 5 minutes to about 1.4 minutes).
Will submit a PR for this issue soon.
**Reporter**: [Cheng Lian](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=lian+cheng) / @liancheng
#### Related issues:
- [Use LRU caching for footers in ParquetInputFormat.](https://github.com/apache/parquet-java/issues/1394) (relates to)
- [Cleanup FilteringParquetRowInputFormat](https://issues.apache.org/jira/browse/SPARK-2551) (is related to)
- [Reading Parquet InputSplits dominates query execution time when reading off S3](https://issues.apache.org/jira/browse/SPARK-2119) (is related to)
- [Improve Parquet IO Performance within cloud datalakes](https://github.com/apache/parquet-java/issues/2912) (is depended upon by)
**Note**: *This issue was originally created as [PARQUET-16](https://issues.apache.org/jira/browse/PARQUET-16). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
贡献指南
这个仓库没有索引到贡献指南
调研方向
从 parquet-hadoop/src/main/java/parquet/hadoop/ParquetInputFormat.java 中的 getSplits,以及 issue 中链接的 footer 获取路径开始。跟踪 FileStatus 对象的获取方式,然后验证 getSplits 不再为每个 part-file 依次获取它们,并比较报告中描述的 S3 或大文件耗时。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- java
- 领域
- performance
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 35/100