NVIDIA / NVIDIA/TransformerEngine

Linear layer reset_parameters() changes bias zero init to random init near 0

Open
#2,529 2 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
3.5k
Forks
831
Avg merge
3d 11h
Merged PRs (30d)
65

Description

Describe the bug

Linear.bias is initialized to a zero vector. After a call to reset_parameters() rather than setting state to what it should be at init time, it gets set to random init. This is the same flavor of bug as #2528 but much less serious since the mean is still the same, and at least with typical initializations the standard deviation is near zero. It would be cleaner to do it correctly though.

Steps/Code to reproduce bug

docker run --gpus=all -it nvcr.io/nvidia/pytorch:25.11-py3 bash
ipython
In [1]: import transformer_engine as te
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:64: FutureWarning: The pynvml package is deprecated. Please install nvidia-ml-py instead. If you did not install pynvml directly, please report this to the maintainers of the package that installed pynvml for you.
  import pynvml  # type: ignore[import]

In [2]: tel = te.pytorch.Linear(3,2)

In [3]: tel.bias
Out[3]: 
Parameter containing:
tensor([0., 0.], device='cuda:0', requires_grad=True)

In [4]: tel.weight
Out[4]: 
Parameter containing:
tensor([[-0.0293, -0.0364,  0.0211],
        [-0.0061, -0.0232,  0.0171]], device='cuda:0', requires_grad=True)

In [5]: tel.reset_parameters()

In [6]: tel.bias
Out[6]: 
Parameter containing:
tensor([-0.0540,  0.0662], device='cuda:0', requires_grad=True)

Expected behavior

bias should be reset to a zero vector after reset_parameters().

Environment overview (please complete the following information)

  • Environment location: Docker
  • Method of Transformer Engine install: pre-installed in nvidia pytorch docker image
  • If method of install is [Docker], provide docker pull & docker run commands used
    (see steps to repro bug)

Environment details

If NVIDIA docker image is used you don't need to specify these.
Otherwise, please provide:

  • OS version
  • PyTorch version
  • Python version
  • Transformer Engine version
  • CUDA version
  • CUDNN version

Device details

  • GPU model

Additional context

Add any other context about the problem here.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the te.pytorch.Linear entry point and inspect reset_parameters(), reproducing the issue with the Docker and IPython steps in the report. Confirm the initial bias is zero, then verify that calling reset_parameters() leaves it as a zero vector. A regression check should cover this expected behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.