flagos-ai / flagos-ai/FlagScale

Bug of MoE checkpoint.

Open
#1,185 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
537
Forks
174
Avg merge
4d 1h
Merged PRs (30d)
12

Description

Now the parallelism of attention and moe is decouple. So there are so many bugs in convert megatron checkpoint to hf checkpoint, especially the expert_tensor_parallel. Please arrange for someone to handle it.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the Megatron-to-HF checkpoint conversion entry point and tracing how expert_tensor_parallel is handled when attention and MoE parallelism are decoupled. Done means the conversion no longer exhibits the reported MoE checkpoint bugs and produces a usable HF checkpoint; the issue does not name specific files or tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.