NVIDIA / NVIDIA/apex

fusedlayernorm does not work with elementwise_affine=False

Open
#403 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
9k
Forks
1.5k
Avg merge
2d 4h
Merged PRs (30d)
3

Description

in distributed training setting with apex disitributed data parallel, when fused layernorm is used with elementwise_affine=False, doing convolution or elementwise sum or elementwise multiplication after fused layer norm throw following error
RuntimeError: CUDA error: invalid configuration argument
but torch nn layernorm works properly with elementwise_affine=False in same setting
also if i use elementwise_affine=True with fused layer norm in same setting, it works properly.
what is problem?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing fused layernorm under Apex distributed data parallel with elementwise_affine=False, followed by convolution or an elementwise sum or multiplication. Compare the failing case with torch.nn LayerNorm and with elementwise_affine=True, then trace the invalid configuration argument to identify what differs; done means the failing operations work in the reported distributed setting.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.