apache / apache/parquet-java

Parquet Pushdown

未关闭
#2,118 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: Parquet Component: Pig Priority: Major Type: bug
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

Hi,
I am doing some experiments with Apache Parquet to test Predicate pushdown and effect of different row group sizes. My assumptions are:

1) Parquet reader first read the metadata to filter out row groups and data pages
2) Then, it reads only those row groups and data pages which match the filter.
3) The total size of read should be the **sum of row group size and size of meta data**.

I have a wide table with 1184 columns. 2 columns are long type and remaining columns are binary. One of the long column is sorted and unique. I disabled dictionary encoding and compression. My file size is 34GB in CSV. I converted it to Parquet. I tried with two options

1) Generate only 1 File of Parquet (i.e. 43GB)
2) Generate multiple files of Parquet (i.e., overall size 43GB).

I allow only 1 Mapper to eliminate the effect of parallelism.

I have a query to search 1 record from the sorted column. The results are for row group 16MB and data page size of 1MB

When there is only 1 file of Parquet.
Input(s):
Successfully read **1 records (22135659519 bytes)** from: "/output/wide/16777216/1048576"

When there is multiple file of Parquet
Input(s):
Successfully read **1 records (800413428 bytes)** from: "/output/wide/16777216/1048576"

My questions are:
1) Why there is big difference. In one file, I am reading 22GB and with multiple file, It is reading 800MB. This is a bug or what?
2) Why it is not reading 16MB + Size of meta data (which is 252MB). Why it is reading more than that?
3) Can I rely on the pig statistics for estimating bytes read?
4) My assumptions are correct or am I missing something?

Could you please have a look into this problem and guide me if it is a bug ?

Logs are attached with this email.

Thank you

Regards
Rana Faisal

**Environment**: Apache Hadoop 2.7.0
Apache Pig 0.17.0
Apache Parquet 1.9.0
**Reporter**: [Rana Faisal Munir](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=ranafaisal342)
#### Original Issue Attachments:
- [selc_32_16777216_1048576_sorted_multiplefile.log](https://issues.apache.org/jira/secure/attachment/12905699/selc_32_16777216_1048576_sorted_multiplefile.log)
- [selc_32_16777216_1048576_sorted.log](https://issues.apache.org/jira/secure/attachment/12905700/selc_32_16777216_1048576_sorted.log)

**Note**: *This issue was originally created as [PARQUET-1192](https://issues.apache.org/jira/browse/PARQUET-1192). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

首先查看单文件和多文件运行的附加日志,然后在所述的 Hadoop 2.7.0、Pig 0.17.0 和 Parquet 1.9.0 环境中重现该查询。将报告的读取字节数与 row-group 和数据页设置进行比较。完成的标准是解释差异,并确定 statistics 或 predicate pushdown 行为是否表明存在一个已确认的 bug。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
data-engineering
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。