Lightning-AI / Lightning-AI/pytorch-lightning

ReduceLROnPlateu within configure_optimizers behave abnormally

Open
#20,829 3 comments 1 reaction 1 assignee View on GitHub

Nobody has claimed this yet.

bug lr scheduler ver: 2.5.x
Dominant language
Python
Stars
31.4k
Forks
3.8k
Avg merge
6d 7h
Merged PRs (30d)
6

Description

### Bug description

Got error
```[python]
File "c:\Users\sean\miniconda3\envs\keras+torch+pl\Lib\site-packages\lightning\pytorch\loops\training_epoch_loop.py", line 459, in _update_learning_rates
raise MisconfigurationException(
lightning.fabric.utilities.exceptions.MisconfigurationException: ReduceLROnPlateau conditioned on metric val/loss which is not available. Available metrics are: ['lr-AdamW/pg1', 'lr-AdamW/pg2', 'train/a_pcc', 'train/loss']. Condition can be set using `monitor` key in lr scheduler dict
```

Here is the `configure_optimizers` function:
```[python]
@final
def configure_optimizers(self):

decay, no_decay = [], []
for name, param in self.named_parameters():
if not param.requires_grad:
continue
if "bias" in name or "Norm" in name:
no_decay.append(param)
else:
decay.append(param)

grouped_params = [
{"params": decay, "weight_decay": self.weight_decay, "lr": self.lr * 0.3},
{
"params": no_decay,
"weight_decay": self.weight_decay,
"lr": self.lr * 1.7,
},
]

optimizer = self.optmizer_class(
grouped_params, lr=self.lr, weight_decay=self.weight_decay
)

scheduler = self.lr_scheduler_class(
optimizer, **self.lr_scheduler_args if self.lr_scheduler_args else {}
)
scheduler = {
"scheduler": self.lr_scheduler_class(
optimizer, **self.lr_scheduler_args if self.lr_scheduler_args else {}
),
"monitor": "val/loss",
"interval": "epoch",
"frequency": 1,
# "strict": False,
}
return {"optimizer": optimizer, "lr_scheduler": scheduler}
```

The `lr_scheduler_class` is passed in as
```[yaml]
lr_scheduler_class: torch.optim.lr_scheduler.ReduceLROnPlateau
lr_scheduler_args:
mode: min
factor: 0.5
patience: 10
threshold: 0.0001
threshold_mode: rel
cooldown: 5
min_lr: 1.e-9
eps: 1.e-08
```
(using yaml and CLI, which, I think, is not the case here)

It seems that I got the error at the end of the training epoch, as I just see the progress bar reports train/loss. The validation epoch is not finished, but the scheduler is called.

I am quite sure that val/loss is available after validation epoch is finished, because progress bar can correctly display it.

### What version are you seeing the problem on?

v2.5

### Reproduced in studio

_No response_

### How to reproduce the bug

```python

```

### Error messages and logs

```
# Error messages and logs here please
```

### Environment

StatusCode : 200
StatusDescription : OK
Content : # Copyright The Lightning AI team.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the...
RawContent : HTTP/1.1 200 OK
Connection: keep-alive
Content-Security-Policy: default-src 'none'; style-src 'unsafe-inline'; sandbox
Strict-Transport-Security: max-age=31536000
X-Content-Type-Options: nosniff
...
Forms : {}
Headers : {[Connection, keep-alive], [Content-Security-Policy, default-src 'none'; style-src 'unsafe-inline'; sandbox], [Strict-Transport-Security, max-age=31536000],
[X-Content-Type-Options, nosniff]...}
Images : {}
InputFields : {}
Links : {}
ParsedHtml : mshtml.HTMLDocumentClass
RawContentLength : 2775

### More info

_No response_

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.