flagos-ai / flagos-ai/FlagScale
Bug of MoE checkpoint.
Open
- Dominant language
- Python
- Stars
- 537
- Forks
- 174
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 12
Description
Now the parallelism of attention and moe is decouple. So there are so many bugs in convert megatron checkpoint to hf checkpoint, especially the expert_tensor_parallel. Please arrange for someone to handle it.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the Megatron-to-HF checkpoint conversion entry point and tracing how expert_tensor_parallel is handled when attention and MoE parallelism are decoupled. Done means the conversion no longer exhibits the reported MoE checkpoint bugs and produces a usable HF checkpoint; the issue does not name specific files or tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100