huggingface / huggingface/datasets

Datasets.map is severely broken

Open
#6,319 15 comments 8 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Describe the bug

Regardless of how many cores I used, I have 16 or 32 threads, map slows down to a crawl at around 80% done, lingers maybe until 97% extremely slowly and NEVER finishes the job. It just hangs.

After watching this for 27 hours I control-C out of it. Until the end one process appears to be doing something, but it never ends.

I saw some comments about fast tokenizers using Rust and all and tried different variations. NOTHING works.

### Steps to reproduce the bug

Running it without breaking the dataset into parts results in the same behavior. The loop was an attempt to see if this was a RAM issue.

for idx in range(100):
dataset = load_dataset("togethercomputer/RedPajama-Data-1T-Sample", cache_dir=cache_dir, split=f'train[{idx}%:{idx+1}%]')
dataset = dataset.map(partial(tokenize_fn, tokenizer), batched=False, num_proc=1, remove_columns=["text", "meta"])
dataset.save_to_disk(training_args.cache_dir + f"/training_data_{idx}")

### Expected behavior

I expect map to run at more or less the same speed it starts with and FINISH its processing.

### Environment info

Python 3.8, same with 3.10 makes no difference.
Ubuntu 20.04,

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.