deepspeedai / deepspeedai/DeepSpeed
Regarding train_batch_size
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
According to the documentation,
train_batch_size is aggregated by the batch size that a single GPU processes in one forward/backward pass (a.k.a., train_step_batch_size), the gradient accumulation steps (a.k.a., gradient_accumulation_steps), and the number of GPUs.
How exactly are these variables related? If no gradient accumulation is used, I believe its just train_batch_size = train_step_batch_size * n_gpu, but how does gradient accumulation come into play? I get a RecursionError: maximum recursion depth exceeded while calling a Python object when the batch sizes in my script and the deepspeed config disagree. Here is the complete error:
> finished creating GPT2 datasets ...
setting training data start iteration to 0
setting validation data start iteration to 0
done with setups ...
time (ms) | model and optimizer: 2350.12 | train/valid/test data iterators: 1179.34
training ...
Traceback (most recent call last):
File "pretrain_gpt2.py", line 156, in <module>
pretrain(train_valid_test_datasets_provider, model_provider, forward_step,
File "/root/megatron-3d/megatron/training.py", line 97, in pretrain
iteration = train(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 481, in train
loss_dict, skipped_iter = train_step(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 324, in train_step
return train_step_pipe(model, data_iterator)
File "/root/megatron-3d/megatron/training.py", line 358, in train_step_pipe
loss = model.train_batch(data_iter=data_iterator)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 273, in train_batch
self._exec_schedule(sched)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 1162, in _exec_schedule
self._exec_instr(**cmd.kwargs)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 621, in _exec_load_micro_batch
batch = self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
[Previous line repeated 978 more times]
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 469, in _next_batch
batch = self.batch_fn(batch)
File "pretrain_gpt2.py", line 110, in get_batch_pipe
return fp32_to_fp16((tokens, position_ids, attention_mask)), fp32_to_fp16((labels, loss_mask))
File "/root/megatron-3d/megatron/fp16/fp16.py", line 53, in fp32_to_fp16
return conversion_helper(val, half_conversion)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in conversion_helper
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in <listcomp>
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 37, in conversion_helper
return conversion(val)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 48, in half_conversion
if isinstance(val_typecheck, (Parameter, Variable)):
File "/root/anaconda3/lib/python3.8/site-packages/torch/autograd/variable.py", line 7, in __instancecheck__
return isinstance(other, torch.Tensor)
RecursionError: maximum recursion depth exceeded while calling a Python object
Traceback (most recent call last):
File "pretrain_gpt2.py", line 156, in <module>
pretrain(train_valid_test_datasets_provider, model_provider, forward_step,
File "/root/megatron-3d/megatron/training.py", line 97, in pretrain
iteration = train(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 481, in train
loss_dict, skipped_iter = train_step(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 324, in train_step
return train_step_pipe(model, data_iterator)
File "/root/megatron-3d/megatron/training.py", line 358, in train_step_pipe
loss = model.train_batch(data_iter=data_iterator)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 273, in train_batch
self._exec_schedule(sched)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 1162, in _exec_schedule
self._exec_instr(**cmd.kwargs)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 621, in _exec_load_micro_batch
batch = self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
[Previous line repeated 978 more times]
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 469, in _next_batch
batch = self.batch_fn(batch)
File "pretrain_gpt2.py", line 110, in get_batch_pipe
return fp32_to_fp16((tokens, position_ids, attention_mask)), fp32_to_fp16((labels, loss_mask))
File "/root/megatron-3d/megatron/fp16/fp16.py", line 53, in fp32_to_fp16
return conversion_helper(val, half_conversion)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in conversion_helper
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in <listcomp>
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 37, in conversion_helper
return conversion(val)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 48, in half_conversion
if isinstance(val_typecheck, (Parameter, Variable)):
File "/root/anaconda3/lib/python3.8/site-packages/torch/autograd/variable.py", line 7, in __instancecheck__
return isinstance(other, torch.Tensor)
RecursionError: maximum recursion depth exceeded while calling a Python object
Traceback (most recent call last):
File "pretrain_gpt2.py", line 156, in <module>
pretrain(train_valid_test_datasets_provider, model_provider, forward_step,
File "/root/megatron-3d/megatron/training.py", line 97, in pretrain
iteration = train(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 481, in train
loss_dict, skipped_iter = train_step(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 324, in train_step
return train_step_pipe(model, data_iterator)
File "/root/megatron-3d/megatron/training.py", line 358, in train_step_pipe
loss = model.train_batch(data_iter=data_iterator)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 273, in train_batch
self._exec_schedule(sched)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 1162, in _exec_schedule
self._exec_instr(**cmd.kwargs)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 621, in _exec_load_micro_batch
batch = self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
[Previous line repeated 978 more times]
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 469, in _next_batch
batch = self.batch_fn(batch)
File "pretrain_gpt2.py", line 110, in get_batch_pipe
return fp32_to_fp16((tokens, position_ids, attention_mask)), fp32_to_fp16((labels, loss_mask))
File "/root/megatron-3d/megatron/fp16/fp16.py", line 53, in fp32_to_fp16
return conversion_helper(val, half_conversion)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in conversion_helper
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in <listcomp>
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 37, in conversion_helper
return conversion(val)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 48, in half_conversion
if isinstance(val_typecheck, (Parameter, Variable)):
File "/root/anaconda3/lib/python3.8/site-packages/torch/autograd/variable.py", line 7, in __instancecheck__
return isinstance(other, torch.Tensor)
RecursionError: maximum recursion depth exceeded while calling a Python object
Traceback (most recent call last):
File "pretrain_gpt2.py", line 156, in <module>
pretrain(train_valid_test_datasets_provider, model_provider, forward_step,
File "/root/megatron-3d/megatron/training.py", line 97, in pretrain
iteration = train(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 481, in train
loss_dict, skipped_iter = train_step(forward_step_func,
File "/root/megatron-3d/megatron/training.py", line 324, in train_step
return train_step_pipe(model, data_iterator)
File "/root/megatron-3d/megatron/training.py", line 358, in train_step_pipe
loss = model.train_batch(data_iter=data_iterator)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 273, in train_batch
self._exec_schedule(sched)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 1162, in _exec_schedule
self._exec_instr(**cmd.kwargs)
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 621, in _exec_load_micro_batch
batch = self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 480, in _next_batch
return self._next_batch()
[Previous line repeated 978 more times]
File "/root/anaconda3/lib/python3.8/site-packages/deepspeed/runtime/pipe/engine.py", line 469, in _next_batch
batch = self.batch_fn(batch)
File "pretrain_gpt2.py", line 110, in get_batch_pipe
return fp32_to_fp16((tokens, position_ids, attention_mask)), fp32_to_fp16((labels, loss_mask))
File "/root/megatron-3d/megatron/fp16/fp16.py", line 53, in fp32_to_fp16
return conversion_helper(val, half_conversion)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in conversion_helper
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 38, in <listcomp>
rtn = [conversion_helper(v, conversion) for v in val]
File "/root/megatron-3d/megatron/fp16/fp16.py", line 37, in conversion_helper
return conversion(val)
File "/root/megatron-3d/megatron/fp16/fp16.py", line 48, in half_conversion
if isinstance(val_typecheck, (Parameter, Variable)):
File "/root/anaconda3/lib/python3.8/site-packages/torch/autograd/variable.py", line 7, in __instancecheck__
return isinstance(other, torch.Tensor)
RecursionError: maximum recursion depth exceeded while calling a Python object
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with pretrain_gpt2.py at get_batch_pipe, then trace train_step_pipe in megatron/training.py into DeepSpeed's pipeline engine _next_batch and the conversion_helper path in megatron/fp16/fp16.py. Clarify how train_batch_size, train_step_batch_size, gradient_accumulation_steps, and GPU count relate, and document the expected behavior when configuration values disagree.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100