mindspore-ai / mindspore-ai/hyper-parallel

[Bug]在fully_shard接口传入replicate_params时,第一个step grad_norm异常比baseline大。

Open
#303 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
53
Forks
63
Avg merge
23h 45m
Merged PRs (30d)
63

Description

该问题是怎么引起的?

网络中部分参数采取不切,将这部分参数传给replicate_params,拉起训练,第一个step grad nrom较大。其他参数配置不改

重现步骤
报错信息

不报错,grad norm异常,baseline grad norm是28.97633362。传replicate_params参数时训练的grad norm是48.47068405

schema_version: 1
source: gitcode
gitcode_repo: mindspore/hyper-parallel
gitcode_issue: 129
source_url: https://gitcode.com/mindspore/hyper-parallel/issues/129

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing training with fully_shard and replicate_params, comparing the first-step grad norm with the baseline configuration. No files or tests are named; the issue is done when the replicated-parameter configuration no longer produces the abnormal first-step norm and matches the baseline behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.