apache / apache/parquet-java

Reuse hadoop file status and footer in ParquetRecordReader

Open
#2,862 0 comments 0 reactions 0 assignees View on GitHub
Component: Hadoop Component: Parquet Priority: Major Type: enhancement
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

## DESCRIPTION

Spark will send a listStatus RPC to get hadoop file status and read the parquet file footer before reading the parquet file. And send a same listStatus RPC to get the same hadoop file status and read the footer again in ParquetRecordReader. We can reuse the file status and the footer.
## PLANS

Save the hadoop file status in the ParquetMetadata and save the ParquetMetadata in the input split, so we can reuse them when init a new ParquetRecordReader.

**Reporter**: [Wan Kun](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=wankun)
#### PRs and other links:
- [GitHub Pull Request #1242](https://github.com/apache/parquet-mr/pull/1242)

**Note**: *This issue was originally created as [PARQUET-2415](https://issues.apache.org/jira/browse/PARQUET-2415). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.