NVIDIA / NVIDIA/apex

Import broken for CPU-only machines: AttributeError: 'NoneType' object has no attribute 'split'

Open
#378 3 comments 3 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

I'm trying to run some code that imports apex on my laptop (without a GPU) for debugging purposes, but the import throws an error. Version info and stacktrace below:

mac osx 10.13.6 (high sierra)
python 3.7.3
torch 1.1.0.post2

git clone https://github.com/NVIDIA/apex && cd apex
pip install -v --no-cache-dir ./
$ python3
Python 3.7.3 (default, Apr  9 2019, 13:13:38)
[Clang 10.0.0 (clang-1000.11.45.5)] on darwin
Type "help", "copyright", "credits" or "license" for more information.
>>> import apex
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/Users/kfsilvers/apex/apex/__init__.py", line 5, in <module>
    from . import amp
  File "/Users/kfsilvers/apex/apex/amp/__init__.py", line 1, in <module>
    from .amp import init, half_function, float_function, promote_function,\
  File "/Users/kfsilvers/apex/apex/amp/amp.py", line 3, in <module>
    from .lists import functional_overrides, torch_overrides, tensor_overrides
  File "/Users/kfsilvers/apex/apex/amp/lists/torch_overrides.py", line 77, in <module>
    if utils.get_cuda_version() >= (9, 1, 0):
  File "/Users/kfsilvers/apex/apex/amp/utils.py", line 9, in get_cuda_version
    return tuple(int(x) for x in torch.version.cuda.split('.'))
AttributeError: 'NoneType' object has no attribute 'split'

Seems like the case where torch.version.cuda returns None isn't handled properly in two places:

  1. apex/amp/utils.py
 def get_cuda_version():
     return tuple(int(x) for x in torch.version.cuda.split('.'))

Seems like we could change this to:

def get_cuda_version():
    if torch.version.cuda is not None:
        return tuple(int(x) for x in torch.version.cuda.split('.'))
    else:
        return None
  1. apex/amp/lists/torch_overrides.py
if utils.get_cuda_version() >= (9, 1, 0):
    FP16_FUNCS.extend(_bmms)
else:
    FP32_FUNCS.extend(_bmms)

We'll also need to add a check for None here:

cuda_version = utils.get_cuda_version()
if cuda_version is not None:
    if cuda_version >= (9, 1, 0):
        FP16_FUNCS.extend(_bmms)
    else:
        FP32_FUNCS.extend(_bmms)

Here's where I confess I know nothing about apex. Should there still be a call to FP32_FUNCS.extend(_bmms) in the case where cuda_version is None i.e. there is no GPU available, or no?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the CPU-only reproduction in the issue, then read apex/amp/utils.py and apex/amp/lists/torch_overrides.py around the reported calls. Determine the expected behavior when torch.version.cuda is None and verify the import path on a CPU-only PyTorch install. Done means importing apex no longer raises this AttributeError and the affected override handling is covered.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.