huggingface / huggingface/datasets

IterableDataset: split by node and map may preprocess samples that will be skipped anyway

Open
#5,961 9 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

There are two ways an iterable dataset can be split by node:
1. if the number of shards is a factor of number of GPUs: in that case the shards are evenly distributed per GPU
2. otherwise, each GPU iterate on the data and at the end keeps 1 sample out of n(GPUs) - skipping the others.

In case 2. it's therefore possible to have the same examples passed to `prepare_dataset` for each GPU.

This doesn't sound optimized though, because it runs the preprocessing on samples that won't be used in the end.

Could you open a new issue so that we can discuss about this and find a solution ?

_Originally posted by @lhoestq in https://github.com/huggingface/datasets/issues/5360#issuecomment-1592729051_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.