huggingface / huggingface/datasets

Arrow map type in parquet files unsupported

Open
#5,612 4 comments 5 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Describe the bug

When I try to load parquet files that were processed with Spark, I get the following issue:

`ValueError: Arrow type map does not have a datasets dtype equivalent.`

Strangely, loading the dataset with `streaming=True` solves the issue.

### Steps to reproduce the bug

The dataset is private, but this can be reproduced with any dataset that has Arrow maps.

### Expected behavior

Loading the dataset no matter whether streaming is True or not.

### Environment info

- `datasets` version: 2.10.1
- Platform: Linux-5.15.0-1029-gcp-x86_64-with-glibc2.31
- Python version: 3.10.7
- PyArrow version: 8.0.0
- Pandas version: 1.4.2

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.