deepspeedai / deepspeedai/DeepSpeed
Out of memory error in the middle of training
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
I am running essentially a fine-tuning script on GPT-Neo here. The issue is that with run command like
PYTHONPATH=~/DeepSpeed deepspeed --num_gpus 4 distill.py --deepspeed_config ds_config_zero3.json
At a certain point (around 1000 iters) it OOMs, even with zero-3 offloading. Here are the wandb logs for both runs. I am using @samyam 's branch full-precision-for-stage3 because the model was trained in bf16, hence breaks with fp16 fine tuning.
Weirdly enough, multi-GPU runs are slower than single-GPU ones, though that too OOMed at the same step.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the reported command with distill.py and ds_config_zero3.json, then compare the linked Weights & Biases logs for single- and multi-GPU runs. Check what changes near iteration 1000 and determine whether the OOM and slowdown can be reproduced; done means the cause is identified and a validated resolution or clear diagnosis is recorded.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100