deepspeedai / deepspeedai/DeepSpeedExamples

Actor loss nan and Resizing model embedding

Open
#922 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
6.8k
Forks
1.1k
Avg merge
2d 16h
Merged PRs (30d)
1

Description

The model I use is GPT-2 124M. When resizing model embeddings during the training of STF and RW, I often encounter issues where the generated answers consist entirely of zeros. This causes both the log probabilities and actor loss to become NaN (Not a Number). I have noticed that resizing the embeddings can lead to the generation of token IDs that exceed the vocabulary size. I suspect this may be contributing to the problem. However, when I don't resize the model's embeddings and train STF and RW, I do not experience this issue during RLHF training. I don't know why.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the GPT-2 124M training case with embedding resizing enabled for STF and RW, then inspect generated token IDs, vocabulary bounds, zero-only outputs, log probabilities, and actor loss. Compare the same training without resizing. Done means the resizing-related cause is identified and the training behavior is verified without invalid token IDs or NaN losses.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.