apache / apache/parquet-java

ParquetReader directory read using InputFile

未关闭
#2,856 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: Hadoop Component: Java Component: Parquet Priority: Major Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

Using Hadoop Path, users can read from directories (excluding hidden files) using a single ParquetReader. This is not yet possible using the InputFile API. To resolve this, a generic way should be added for directories to be read, consistent with the behaviour of the read API when using a Hadoop Path.

**Reporter**: [Atour Mousavi Gourabi](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=amousavigourabi) / @amousavigourabi

**Note**: *This issue was originally created as [PARQUET-2404](https://issues.apache.org/jira/browse/PARQUET-2404). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

首先,将现有的 Hadoop Path 目录读取行为与 issue 中描述的 InputFile API 进行比较。定义通用目录读取应如何工作,包括排除隐藏文件并保留单个 reader 的行为,然后根据现有 read API 的行为验证结果。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data-engineering
Issue 类型
功能
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。