NVIDIA / NVIDIA/apex

RuntimeError: expected scalar type Float but found Half

Open
#965 7 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

def __init__(self, chi, cho):
    super(DeformConv, self).__init__()
    self.actf = nn.Sequential(
        nn.BatchNorm2d(cho, momentum=BN_MOMENTUM),
        nn.ReLU(inplace=True)
    )
    self.conv = DCN(chi, cho, kernel_size=(3, 3), stride=1, padding=1, dilation=1, deformable_groups=1)
def forward(self, x):
    x = self.conv(x)
    x = self.actf(x)
    return x

My model exsits a DCN module which compiled by c++. when I use amp.initialize(model, optimizer, opt_level="O1"), RuntimeError has happened(expected scalar type Float but found Half) in x=self.conv(x). I try to use x=self.conv(x.float()) to convert type, but not useful.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the shown DeformConv.forward call and inspect how the C++-compiled DCN module handles the tensor dtype under amp.initialize with opt_level="O1". Compare the input and extension expectations around x=self.conv(x), including the attempted float conversion. Done means the reported RuntimeError is resolved for this mixed-precision path; no specific files or tests are named.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.