mindspore-ai / mindspore-ai/hyper-parallel

支持非融合 FSDP 逐参数混合 dtype

Open
#225 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
53
Forks
63
Avg merge
23h 45m
Merged PRs (30d)
63

Description

当前 MindSpore FSDP 非融合逐参数通信路径仍复用 state 级统一 dtype 元数据,混合原始参数 dtype 时会触发统一 dtype 约束或导致错误 cast。需要改为按参数自身的 orig_dtype/reduce_dtype 进行梯度通信与回写,并保留 comm_fusion 路径的统一 dtype 约束。

schema_version: 1
source: gitcode
gitcode_repo: mindspore/hyper-parallel
gitcode_issue: 285
source_url: https://gitcode.com/mindspore/hyper-parallel/issues/285

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the non-fused, per-parameter FSDP gradient communication and write-back path, then trace where state-level dtype metadata is applied. Check mixed original parameter dtypes and verify that each parameter uses its own orig_dtype/reduce_dtype, while the comm_fusion path retains its uniform dtype constraint.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.