Lightning-AI / Lightning-AI/pytorch-lightning

Returning None from training_step with multi GPU DDP training

Open
#5,243 26 comments 4 reactions 1 assignee Claimed by @Borda View on GitHub
distributed feature help wanted priority: 1
Dominant language
Python
Stars
31.4k
Forks
3.8k
Avg merge
6d 7h
Merged PRs (30d)
6

Description

## 🐛 Bug

Returning None from training_step with multi GPU DDP training freezes the training without exception

### To Reproduce
Starting multi-gpu training with a None-returning training_step function.

Example training_step function:
```
def training_step(self, batch, batch_idx):
data, target = batch
model_outputs = self.forward(images)
loss = calc_loss(model_outputs, target)

if torch.isnan(loss) or random.random() < .05:
return None

return loss
```
Example trainer:
```
trainer = Trainer(
gpus=2,
distributed_backend="ddp"
)
```

### Expected behavior

To continue training with skipping the current batch as pointed out at [here](https://pytorch-lightning.readthedocs.io/en/latest/lightning_module.html#training-step).
### Environment

No specific environment is needed to reproduce this bug.

### Additional context

This issue was mentioned here: #4956 but not with specifics.

**Note: By the time this issue being investigated, a help for a workaround would be great!**

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.