ParquetRecordReader support building ParquetFileReader using incoming Footer
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 33
Description
SPARK-33449 discusses adding file meta cache mechanism for Spark SQL to reduce data reading and speed up query.
For the scenarios that need to use ParquetRecordReader, we hope to add an interface to support passing an existing footer to ParquetRecordReader and use this instance to build ParquetFileReader.
**Reporter**: [Yang Jie](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=LuciferYang) / @LuciferYang
**Note**: *This issue was originally created as [PARQUET-1965](https://issues.apache.org/jira/browse/PARQUET-1965). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating ParquetRecordReader and the code that builds ParquetFileReader, then inspect existing reader tests and usages. Implement support for supplying an existing footer and verify that the reader uses it when constructing the ParquetFileReader without changing the current path for callers that do not provide one.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100