mindspore-ai / mindspore-ai/hyper-parallel

[fully_shard]处理chunk_loss时,报错unsharded_param.grad is none

Open
#324 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
53
Forks
63
Avg merge
23h 45m
Merged PRs (30d)
63

Description

该问题是怎么引起的?

fully_shard处理chunk_loss时,报错unsharded_param.grad is none,原因在于在某些求导场景中,dw会延迟到最后一个dx获取后才准备好

重现步骤
报错信息

schema_version: 1
source: gitcode
gitcode_repo: mindspore/hyper-parallel
gitcode_issue: 65
source_url: https://gitcode.com/mindspore/hyper-parallel/issues/65

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the fully_shard handling for chunk_loss and reproduce the reported unsharded_param.grad is none failure. Trace when dw becomes ready relative to the final dx; done when the affected differentiation scenario completes without the missing-gradient error and the reproduction is covered by a regression test.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.