huggingface / huggingface/datasets
Support of num_workers (multiprocessing) in map for IterableDataset
Open
enhancement
- Dominant language
- Python
- Stars
- 22k
- Forks
- 3.4k
- Avg merge
- 5d 7h
- Merged PRs (30d)
- 17
Description
### Feature request
Currently, IterableDataset doesn't support setting num_worker in .map(), which results in slow processing here. Could we add support for it? As .map() can be run in the batch fashion (e.g., batch_size is default to 1000 in datasets), it seems to be doable for IterableDataset as the regular Dataset.
### Motivation
Improving data processing efficiency
### Your contribution
Testing
Contributor guide
Assessment
This issue has not been assessed yet.