deepspeedai / deepspeedai/DeepSpeed
[REQUEST] Can deepspeed initialize multiple datasets/dataloaders?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
When I train a model, I want to evaluate the model and save ckpt at each epoch, enabling the same DP for the training dataloader and eval dataloader.
Now: DeepSpeed can only initialize the training dataset. If using the torch dataloader within the epoch, it will run duplicated.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing DeepSpeed's current training-dataset initialization and how a torch DataLoader is consumed within an epoch. Determine the design needed to initialize separate training and evaluation dataloaders with the same data parallelism, without duplicated evaluation iteration. Done means epoch evaluation and checkpointing can use the separate loader correctly.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100