huggingface / huggingface/datasets

Almost identical datasets, huge performance difference

Open
#5,669 7 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Describe the bug

I am struggling to understand (huge) performance difference between two datasets that are almost identical.

### Steps to reproduce the bug

# Fast (normal) dataset speed:
```python
import cv2
from datasets import load_dataset
from torch.utils.data import DataLoader

dataset = load_dataset("beans", split="train")

for x in DataLoader(dataset.with_format("torch"), batch_size=16, shuffle=True, num_workers=8):
pass
```

The above pass over the dataset takes about 1.5 seconds on my computer.
However, if I re-create (almost) the same dataset, the sweep takes HUGE amount of time: 15 minutes. Steps to reproduce:

```python
def transform(example):
example["image2"] = cv2.imread(example["image_file_path"])
return example

dataset2 = dataset.map(transform, remove_columns=["image"])

for x in DataLoader(dataset2.with_format("torch"), batch_size=16, shuffle=True, num_workers=8):
pass

```

### Expected behavior

Same timings

### Environment info

python==3.10.9
datasets==2.10.1

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.