NVIDIA / NVIDIA/apex

After installing apex, failing to train the original model with 32-bit

Open
#568 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

Hi,

It works when I use apex to train a model with bs=32. However, when bs is set to 16, I fail to train the model without apex, that is, the loss is no longer reduced by using 32 bits, which is originally effective .

Thanks for any suggestion.

Env:
2080Ti
Cuda 10.0
pytorch 1.2

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No source file, test, model, or reproduction command is named. Start by reproducing the batch-size 16 failure in the stated PyTorch 1.2, CUDA 10.0, and 2080Ti environment, comparing training with and without apex; done requires identifying the cause and confirming that 32-bit training reduces the loss as expected.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.