NVIDIA / NVIDIA/apex

Model parallel: an illegal memory access was encountered

Open
#371 30 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

I'm running my model to process really large 3D volumes, so I have to define a model parallel like this:
Class model(....):
def forward(self, x):
#x is on 'cuda:0'
XA=self.A(x)
XB=self.B(XA.to('cuda:1'))
return XB
It runs well using float32, but still I want larger volume size or more channels, so I tried apex and it reported:
RuntimeError: CUDA error: an illegal memory access was encountered
from the scaler.py: self._has_overflow = self._overflow_buf.item()
Any ideas?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the failure reported at scaler.py, specifically self._overflow_buf.item(), and reproduce the described model-parallel run with mixed precision versus float32 on the two CUDA devices. Trace whether the illegal access originates in overflow handling or the preceding model-parallel operation; done means the cause and a verified resolution for this configuration are established.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.