huggingface / huggingface/datasets

Stuck in "Resolving data files..."

Open
#6,359 5 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Describe the bug

I have an image dataset with 300k images, the size of image is 768 * 768.

When I run `dataset = load_dataset("imagefolder", data_dir="/path/to/img_dir", split='train')` in second time, it takes 50 minutes to finish "Resolving data files" part, what's going on in this part?

From my understand, after Arrow files been created in the first run, the second run should not take time longer than one or two minutes.

### Steps to reproduce the bug

# Run following code two times
dataset = load_dataset("imagefolder", data_dir="/path/to/img_dir", split='train')

### Expected behavior

Fast dataset building

### Environment info

- `datasets` version: 2.14.5
- Platform: Linux-5.15.0-60-generic-x86_64-with-glibc2.35
- Python version: 3.10.11
- Huggingface_hub version: 0.17.3
- PyArrow version: 10.0.1
- Pandas version: 1.5.3

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.