deepspeedai / deepspeedai/DeepSpeed

[BUG] Memory leak during autograd backwards

Open
#5,198 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug training
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

Describe the bug
@exnx discovered that using larger gradient_accumulation_steps sizes with GPT-NeoX causes a large memory increase from smaller GAS sizes. When @Quentin-Anthony and I used the PyTorch memory profiler with DS, we saw that buffers allocated during the backwards pass by torch.autograd hang until the end of the training step. These buffers should not be necessary to allocate until the end of the training step. These cause DS to use much more memory than it needs to for larger batch sizes. Using PyTorch 2.0.1.

To Reproduce
I used GPT-NeoX, and ran with config:

# GPT-2 pretraining setup
{
   # parallelism settings ( you will want to change these based on your cluster setup, ideally scheduling pipeline stages
   # across the node boundaries )
   "pipe_parallel_size": 2,
   "model_parallel_size": 1,

   # model settings
   "num_layers": 24,
   "hidden_size": 2048,
   "num_attention_heads": 16,
   "seq_length": 2048,
   "max_position_embeddings": 2048,
   "norm": "layernorm",
   "pos_emb": "rotary",
   "no_weight_tying": true,
   "gpt_j_residual": false,
   "output_layer_parallelism": "column",

   # these should provide some speedup but takes a while to build, set to true if desired
   "scaled_upper_triang_masked_softmax_fusion": false,
   "bias_gelu_fusion": false,
   "rope_fusion": false,
   "layernorm_fusion": false,

   # init methods
   "init_method": "small_init",
   "output_layer_init_method": "wang_init",

   # optimizer settings
   "optimizer": {
     "type": "Adam",
     "params": {
       "lr": 0.0002,
       "betas": [0.9, 0.95],
       "eps":  1.0e-8,
     }
   },
   "min_lr": 0.00002,

   # for all zero_optimization options, see https://www.deepspeed.ai/docs/config-json/#zero-optimizations-for-fp16-training
   "zero_optimization": {
    "stage": 1,
    "allgather_partitions": True,
    "allgather_bucket_size": 500000000,
    "overlap_comm": True,
    "reduce_scatter": True,
    "reduce_bucket_size": 500000000,
    "contiguous_gradients": True,
  },

   # batch / data settings
   "train_micro_batch_size_per_gpu": 1,
   "gradient_accumulation_steps": 4,
   "data_impl": "mmap",

   # activation checkpointing
   "checkpoint_activations": true,
   "checkpoint_num_layers": 1,
   "partition_activations": true,
   "synchronize_each_layer": true,

   # regularization
   "gradient_clipping": 1.0,
   "weight_decay": 0.1,
   "hidden_dropout": 0,
   "attention_dropout": 0,

   # precision settings
   "fp16": {
     "fp16": true,
     "enabled": true,
     "loss_scale": 0,
     "loss_scale_window": 1000,
     "hysteresis": 2,
     "min_loss_scale": 1
   },

   # misc. training settings
   "train_iters": 100,
   "lr_decay_iters": 320000,
   "distributed_backend": "nccl",
   "lr_decay_style": "cosine",
   "warmup": 0.01,
   "checkpoint_factor": 10000,
   "eval_interval": 1000,
   "eval_iters": 10,

   # logging
   "log_interval": 100,
   "steps_per_print": 10,
   "keep_last_n_checkpoints": 4,
   "wall_clock_breakdown": true,
   "profile": true,
   "profile_step_start": 10,
   "profile_step_stop": 50,
   "memory_profiling": true,
   "memory_profiling_path": "/weka/home-jacob/neox-logs/gas4",
}

Expected behavior
I expect these buffers to be deallocated at the end of their use instead of at the end of the training step.

Screenshots

Screen Shot 2024-02-27 at 1 33 11 AM

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No source file or test is named. Start by reproducing the GPT-NeoX configuration with PyTorch 2.0.1 and DeepSpeed's memory profiler, comparing smaller and larger gradient_accumulation_steps values. Done means identifying why autograd buffers remain allocated through the training step and confirming they are released at the end of their use.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.