PyArrow does not correctly read empty categorical columns from Parquet into Pandas DataFrames
Open
Component: Parquet
Component: Python
Type: bug
- Dominant language
- C++
- Stars
- 17.1k
- Forks
- 4.3k
- Avg merge
- 3d 13h
- Merged PRs (30d)
- 88
Description
### Describe the bug, including details regarding any error messages, version, and platform.
As documented in pandas-dev/pandas#48883, if a DataFrame with an empty categorical column is saved into a Parquet and subsequently loaded using `pyarrow`, the column's dtype reverts to `object`. This issue occurs regardless of what engine was used to save the Parquet, and does not occur when using `fastparquet` to load the file instead.
### Component(s)
Parquet, Python
Contributor guide
Assessment
This issue has not been assessed yet.