NVIDIA / NVIDIA/apex

A way to do 32 bit training and then transition to 16 bit.

Open
#788 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

Like many, I'm having issues with scale overflows, but also general instability during training. Nans Infs etc.

This wasn't the case until I upgraded to the new API, which I generally like a lot better.

In any event, if I can avoid Nans / divide by zero for the first 1000 or so iterations, everything becomes stable.

Is there a way to train using 32 bit for the first 1000 or so iterations and then change to 16 when training is more stable?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names no files, tests, or entry points. Start by reviewing the mixed-precision API and its training configuration, then determine how a 32-bit warm-up could transition to 16-bit after roughly 1,000 iterations. A reproducible case and tests covering the transition and numerical stability would define what is done.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.