apache / apache/parquet-java

Add support for Parquet Input Stream optimisations.

未关闭
#3,386 4 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

This issue tracks adding of a new module [parquet-io] with the goal of adding IO optimisations to the parquet-java repository.

The goal is to have a single place where the following optimisations can be implemented:

* Vectored Reads
* Reading the tail of a file in one request rather than multiple small requests (Avoid the parquet footer dance, multiple requests for the pageIndex)
* Small Parquet files are read in a single request
* Sequential prefetching

[Parquet Java Input Stream Optimisations](https://docs.google.com/document/d/1Xdlh23tmCs-KvzHhY2RuwFYmc3xntUKcmwb8yxEl78Y/edit?usp=sharing): Doc explains the features this will implement.

[Analytics Accelerator for S3](https://docs.google.com/document/d/13shy0RWotwfWC_qQksb95PXdi-vSUCKQyDzjoExQEN0/edit?tab=t.0#heading=h.3lc3p7s26rnw): Doc explains IO optimisations made in the analytics accelerator library.

贡献指南

这个仓库没有索引到贡献指南

调研方向

首先阅读 issue 中链接的 Parquet Java Input Stream Optimisations 文档和 Analytics Accelerator for S3 文档。然后围绕向量化读取、尾部和小文件读取、页面索引访问以及顺序预取,定义拟议 parquet-io 模块的范围。完成的标准是形成一致认可的实现计划,并支持所列出的优化。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data-engineering, performance
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
需要澄清
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。