huggingface / huggingface/datasets

HF Datasets data access is extremely slow even when in memory

Open
#6,104 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Describe the bug

Doing a simple `some_dataset[:10]` can take more than a minute.

Profiling it:
image

`some_dataset` is completely in memory with no disk cache.

This is proving fatal to my usage of HF Datasets. Is there a way I can forgo the arrow format and store the dataset as PyTorch tensors so that `_tensorize` is not needed? And is `_consolidate` supposed to take this long?

It's faster to produce the dataset from scratch than to access it from HF Datasets!

### Steps to reproduce the bug

I have uploaded the dataset that causes this problem [here](https://huggingface.co/datasets/NightMachinery/hf_datasets_bug1).

```python
#!/usr/bin/env python3
import sys
import time
import torch
from datasets import load_dataset

def main(dataset_name):
# Start the timer
start_time = time.time()

# Load the dataset from Hugging Face Hub
dataset = load_dataset(dataset_name)

# Set the dataset format as torch
dataset.set_format(type="torch")

# Perform an identity map
dataset = dataset.map(lambda example: example, batched=True, batch_size=20)

# End the timer
end_time = time.time()

# Print the time taken
print(f"Time taken: {end_time - start_time:.2f} seconds")

if __name__ == "__main__":
dataset_name = "NightMachinery/hf_datasets_bug1"
print(f"dataset_name: {dataset_name}")
main(dataset_name)
```

### Expected behavior

_

### Environment info

- `datasets` version: 2.13.1
- Platform: Linux-5.15.0-76-generic-x86_64-with-glibc2.35
- Python version: 3.10.12
- Huggingface_hub version: 0.16.4
- PyArrow version: 12.0.1
- Pandas version: 2.0.3

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.