Investigate using new Hadoop Vectored IO capability
Open
on-hold
- Dominant language
- Java
- Stars
- 107
- Forks
- 29
- Avg merge
- 19h 46m
- Merged PRs (30d)
- 141
Description
This video describes a new vectored read method in the Hadoop FileSystem API: https://www.youtube.com/watch?v=mRaOSxLoCtM This gives significantly better performance when reading from S3. For Sleeper to exploit this we will need the Parquet library to be updated - see https://issues.apache.org/jira/projects/PARQUET/issues/PARQUET-2171?filter=allopenissues and https://github.com/apache/parquet-mr/pull/1103
Contributor guide
Assessment
This issue has not been assessed yet.