apache / apache/parquet-java

[C++] 1.4.1 library allows creation of parquet file w/NULL values for INT types

未关闭
#2,202 6 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
Component: C++ Component: Java Component: Parquet Priority: Major Type: bug
主要语言
Java
星标
3.1k
派生
1.6k
平均合并
3 天 12 小时
30 天内合并 PR
33

描述

The parquet-cpp v1.4.1 library allows generation of parquet files with NULL values for INT type columns which causes unexpected parsing errors in downstream systems ingesting those files.

e.g.,
`Error parsing the parquet file: UNKNOWN can not be applied to a primitive type`

**Reproduction Steps**

OS: CentOS 7.5.1804
Python: 3.4.8

Prerequisites:
- Install the following packages: `Numpy: 1.14.5`, `Pandas: 0.22.0`, `PyArrow: 0.9.0`

Step 1

Generate the parquet file.

`sample_w_null.csv`

```Java

col1,col2,col3,col4,col5
1,2,,4,5
```

`parquet-1361-repro-1.py`

{code}
#!/usr/bin/python

import numpy as np
import pyarrow as pa
import pyarrow.parquet as pq
import pandas as pd

input_file = 'sample_w_null.csv'
output_file = 'int_unknown.parquet'
p_schema = {'col1': np.int32,
'col2': np.int32,
'col3': np.unicode_,
'col4': np.int32,
'col5': np.int32}

df = pd.read_csv(input_file, dtype=p_schema)
table = pa.Table.from_pandas(df)
pq.write_table(table, output_file)
```Java

+Step 2+

Inspect the metadata of the generated file.

{{parquet-1361-repro-2.py}}

```
#!/usr/bin/python

import pyarrow.parquet as pq

for filename in ['int_unknown.parquet']:
pq_file = pq.ParquetFile(filename)
print(pq_file.metadata)
print(pq_file.schema)
print(pq_file.num_row_groups)
print(pq.read_table(filename, columns=['col1','col2','col3','col4','col5']).to_pandas())
```Java

Results

```

created_by: parquet-cpp version 1.4.1-SNAPSHOT
num_columns: 6
num_rows: 1
num_row_groups: 1
format_version: 1.0
serialized_size: 1434

col1: INT32
col2: INT32
col3: INT32 UNKNOWN
col4: INT32
col5: INT32
__index_level_0__: INT64

1
col1 col2 col3 col4 col5
0 1 2 None 4 5
{code}

**Reporter**: [Ken Terada](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=spherified)
#### Original Issue Attachments:
- [parquet-1361-repro-1.py](https://issues.apache.org/jira/secure/attachment/12933498/parquet-1361-repro-1.py)
- [parquet-1361-repro-2.py](https://issues.apache.org/jira/secure/attachment/12933499/parquet-1361-repro-2.py)
- [sample_w_null.csv](https://issues.apache.org/jira/secure/attachment/12933500/sample_w_null.csv)

**Note**: *This issue was originally created as [PARQUET-1361](https://issues.apache.org/jira/browse/PARQUET-1361). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 parquet-1361-repro-1.py 和 sample_w_null.csv 开始,复现生成的文件,然后使用 parquet-1361-repro-2.py 检查其 schema 和 NULL 处理。追踪负责 INT32 UNKNOWN 列的写入路径,并将结果与预期的 Parquet 元数据进行比较;完成的标准是 nullable INT 列不再产生无效的 UNKNOWN primitive type,并且下游可以读取该复现结果。

由索引模型根据 Issue 内容生成。

评估

技术栈
cpp, pandas, python
领域
data-engineering, databases
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。