tensorflow / tensorflow/probability
NaN Losses for GradientDescentOptimizer and MomentumOptimizer
Nobody has claimed this yet.
- Dominant language
- Jupyter Notebook
- Stars
- 4.4k
- Forks
- 1.1k
- PR merge metrics
- No merged PRs in 30d
Description
In cifar10_bnn.py,:
For fake data, using tf.compat.v1.train.GradientDescentOptimizer and tf.compat.v1.train.MomentumOptimizer quickly leads to NaN losses and KL (after one or two iterations).
It works for very small learning rates but then VERY slowly converge.
Do you have an idea why?
I have to put a learning rate of 10^-9 to make it work. (And it is very sensitive)
Thanks!
Belhal
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with cifar10_bnn.py and reproduce the report using fake data, tf.compat.v1.train.GradientDescentOptimizer, and tf.compat.v1.train.MomentumOptimizer. Track the loss and KL values across the first iterations and compare behavior at the reported learning rates. Done means identifying the source of the NaNs and documenting or correcting the optimizer behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- tensorflow
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100