tensorflow / tensorflow/probability

NaN Losses for GradientDescentOptimizer and MomentumOptimizer

Open
#833 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
4.4k
Forks
1.1k
PR merge metrics
No merged PRs in 30d

Description

In cifar10_bnn.py,:
For fake data, using tf.compat.v1.train.GradientDescentOptimizer and tf.compat.v1.train.MomentumOptimizer quickly leads to NaN losses and KL (after one or two iterations).

It works for very small learning rates but then VERY slowly converge.

Do you have an idea why?
I have to put a learning rate of 10^-9 to make it work. (And it is very sensitive)

Thanks!
Belhal

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with cifar10_bnn.py and reproduce the report using fake data, tf.compat.v1.train.GradientDescentOptimizer, and tf.compat.v1.train.MomentumOptimizer. Track the loss and KL values across the first iterations and compare behavior at the reported learning rates. Done means identifying the source of the NaNs and documenting or correcting the optimizer behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
tensorflow
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.