deepspeedai / deepspeedai/DeepSpeedExamples

use bloom-350m to train reward model in step2

Open
#356 7 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

deespeed chat new-config
Dominant language
Python
Stars
6.8k
Forks
1.1k
Avg merge
2d 16h
Merged PRs (30d)
1

Description

I want to train bloom_350m in chinese dataset, and run run_350m.sh, change the model_name_or_path. But the loss is nan, how should I solve it? Is the argument "num_padding_at_beginning" cause this?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with run_350m.sh and the step2 training path, then reproduce the NaN loss using the Chinese dataset and BLOOM-350M model path. Inspect how num_padding_at_beginning is passed and identify whether it causes the failure; done means a confirmed cause and a documented corrective change or configuration.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, shell
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.