apache / apache/arrow

PyArrow does not correctly read empty categorical columns from Parquet into Pandas DataFrames

Open
#45,192 0 comments 0 reactions 0 assignees View on GitHub
Component: Parquet Component: Python Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the bug, including details regarding any error messages, version, and platform.

As documented in pandas-dev/pandas#48883, if a DataFrame with an empty categorical column is saved into a Parquet and subsequently loaded using `pyarrow`, the column's dtype reverts to `object`. This issue occurs regardless of what engine was used to save the Parquet, and does not occur when using `fastparquet` to load the file instead.

### Component(s)

Parquet, Python

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.