apache / apache/parquet-java

[C++] 1.4.1 library allows creation of parquet file w/NULL values for INT types

Open
#2,202 6 comments 0 reactions 0 assignees View on GitHub
Component: C++ Component: Java Component: Parquet Priority: Major Type: bug
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

The parquet-cpp v1.4.1 library allows generation of parquet files with NULL values for INT type columns which causes unexpected parsing errors in downstream systems ingesting those files.

e.g.,
`Error parsing the parquet file: UNKNOWN can not be applied to a primitive type`

**Reproduction Steps**

OS: CentOS 7.5.1804
Python: 3.4.8

Prerequisites:
- Install the following packages: `Numpy: 1.14.5`, `Pandas: 0.22.0`, `PyArrow: 0.9.0`

Step 1

Generate the parquet file.

`sample_w_null.csv`

```Java

col1,col2,col3,col4,col5
1,2,,4,5
```

`parquet-1361-repro-1.py`

{code}
#!/usr/bin/python

import numpy as np
import pyarrow as pa
import pyarrow.parquet as pq
import pandas as pd

input_file = 'sample_w_null.csv'
output_file = 'int_unknown.parquet'
p_schema = {'col1': np.int32,
'col2': np.int32,
'col3': np.unicode_,
'col4': np.int32,
'col5': np.int32}

df = pd.read_csv(input_file, dtype=p_schema)
table = pa.Table.from_pandas(df)
pq.write_table(table, output_file)
```Java

+Step 2+

Inspect the metadata of the generated file.

{{parquet-1361-repro-2.py}}

```
#!/usr/bin/python

import pyarrow.parquet as pq

for filename in ['int_unknown.parquet']:
pq_file = pq.ParquetFile(filename)
print(pq_file.metadata)
print(pq_file.schema)
print(pq_file.num_row_groups)
print(pq.read_table(filename, columns=['col1','col2','col3','col4','col5']).to_pandas())
```Java

Results

```

created_by: parquet-cpp version 1.4.1-SNAPSHOT
num_columns: 6
num_rows: 1
num_row_groups: 1
format_version: 1.0
serialized_size: 1434

col1: INT32
col2: INT32
col3: INT32 UNKNOWN
col4: INT32
col5: INT32
__index_level_0__: INT64

1
col1 col2 col3 col4 col5
0 1 2 None 4 5
{code}

**Reporter**: [Ken Terada](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=spherified)
#### Original Issue Attachments:
- [parquet-1361-repro-1.py](https://issues.apache.org/jira/secure/attachment/12933498/parquet-1361-repro-1.py)
- [parquet-1361-repro-2.py](https://issues.apache.org/jira/secure/attachment/12933499/parquet-1361-repro-2.py)
- [sample_w_null.csv](https://issues.apache.org/jira/secure/attachment/12933500/sample_w_null.csv)

**Note**: *This issue was originally created as [PARQUET-1361](https://issues.apache.org/jira/browse/PARQUET-1361). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with parquet-1361-repro-1.py and sample_w_null.csv to reproduce the generated file, then use parquet-1361-repro-2.py to inspect its schema and NULL handling. Trace the writer path responsible for the INT32 UNKNOWN column and compare the result with expected Parquet metadata; done means nullable INT columns no longer produce an invalid UNKNOWN primitive type and the reproduction can be read downstream.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, pandas, python
Domain
data-engineering, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.