deepspeedai / deepspeedai/DeepSpeed

[BUG]Zero stage3 can not save model weights correctly!

Open
#3,841 8 comments 5 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug training
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

Describe the bug
Traing the llama-7b model with zero stage3 and set stage3_gather_16bit_weights_on_model_save to true in ds_config.json, but the size of saved pytorch-model.bin is only 610K. It is strange that the saved model in checkpoint is normal.

The deepspeed version is 0.9.6

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing Llama-7b training with ZeRO stage 3 and stage3_gather_16bit_weights_on_model_save enabled in ds_config.json. Compare the saved pytorch-model.bin with the normal checkpoint weights; done means the model save contains correctly gathered 16-bit weights rather than a 610K file.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.