google-research / google-research/augmix
Nan loss for ResNext backbone trained on cifar 100
Open
- Dominant language
- Python
- Stars
- 989
- Forks
- 156
- PR merge metrics
- No merged PRs in 30d
Description
Thank you for your work. While trying your code for the Resnext backbone on cifar100, I get nan values for the training loss. As mentioned in the published paper, I use the initial learning rate of 0.1 for SGD with cosine scheduling.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the ResNext backbone on CIFAR-100 with SGD at an initial learning rate of 0.1 and cosine scheduling, then inspect where the training loss first becomes NaN. Done means the cause is identified and training produces finite loss values under the reported setup.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100