mindspore-ai / mindspore-ai/hyper-parallel

[Bug]在fully_shard接口传入replicate_params时,第一个step grad_norm异常比baseline大。

Open
#735 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
53
Forks
63
Avg merge
23h 45m
Merged PRs (30d)
63

Description

该问题是怎么引起的?

网络中部分参数采取不切,将这部分参数传给replicate_params,拉起训练,第一个step grad nrom较大。其他参数配置不改

重现步骤
报错信息

不报错,grad norm异常,baseline grad norm是28.97633362。传replicate_params参数时训练的grad norm是48.47068405

schema_version: 1
source: gitcode
gitcode_repo: mindspore/hyper-parallel
gitcode_issue: 129
source_url: https://gitcode.com/mindspore/hyper-parallel/issues/129

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the fully_shard interface and reproduce the first training step with and without replicate_params, keeping all other parameters unchanged. Compare the resulting grad norm with the reported baselines of 28.97633362 and 48.47068405, then trace the discrepancy and verify that the first-step value matches the baseline after the fix.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.