apache / apache/parquet-java

Vectorized API to support Column Index in Apache Spark

未关闭
#2,476 10 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: Java Component: Parquet Priority: Major Type: enhancement
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

As per the comment on https://issues.apache.org/jira/browse/SPARK-26345. Its seems like Apache Spark doesn't support Column Index until we disable vectorizedReader in Spark - which will have other performance implications. As per  @zivanfi  , parquet-mr should implement a Vectorized API. Is it already implemented or any pull request for the same?

**Reporter**: [Felix Kizhakkel Jose](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=FelixKJose) / @FelixKJose

**Note**: *This issue was originally created as [PARQUET-1830](https://issues.apache.org/jira/browse/PARQUET-1830). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

该 issue 指向 SPARK-26345,并询问 parquet-mr 是否已经具备向量化 API 或相关的 pull request;首先查看链接的讨论,并检查当前的 parquet-java 实现。由于报告询问的是是否存在支持,而不是指定一项具体更改,因此未定义 Done。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data-engineering
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
停滞
描述清晰度
需要澄清
新手友好度
20/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。