deepspeedai / deepspeedai/DeepSpeed

Out of memory error in the middle of training

Open
#1,003 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

I am running essentially a fine-tuning script on GPT-Neo here. The issue is that with run command like

PYTHONPATH=~/DeepSpeed deepspeed --num_gpus 4 distill.py --deepspeed_config ds_config_zero3.json

At a certain point (around 1000 iters) it OOMs, even with zero-3 offloading. Here are the wandb logs for both runs. I am using @samyam 's branch full-precision-for-stage3 because the model was trained in bf16, hence breaks with fp16 fine tuning.
Weirdly enough, multi-GPU runs are slower than single-GPU ones, though that too OOMed at the same step.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the reported command with distill.py and ds_config_zero3.json, then compare the linked Weights & Biases logs for single- and multi-GPU runs. Check what changes near iteration 1000 and determine whether the OOM and slowdown can be reproduced; done means the cause is identified and a validated resolution or clear diagnosis is recorded.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.