NVIDIA / NVIDIA/apex

fail to use O1 level for AdaptiveLogSoftmaxWithLoss

Open
#556 1 comment 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

File "/opt/conda/envs/python3.6/lib/python3.6/site-packages/torch/nn/modules/adaptive.py", line 186, in forward
output.index_copy_(0, row_indices, local_logprob.squeeze(1))
RuntimeError: Expected object of scalar type Float but got scalar type Half for argument #4 'source'

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with torch/nn/modules/adaptive.py at line 186 and trace the AdaptiveLogSoftmaxWithLoss forward path under Apex O1 mixed precision. Reproduce the reported Float/Half mismatch and determine the expected behavior for this operation. Done means the forward pass completes successfully at O1 without the reported scalar-type error.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.