NVIDIA / NVIDIA/Megatron-LM

allow support for heterogeneous models with MoE

Open
#5,573 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
17.9k
Forks
4.5k
Avg merge
4d 6h
Merged PRs (30d)
271

Description

clear issues with router bias for cases where MoE layers may have different number of experts

Contributor guide

Open the contributing guide

Research direction

No file, test, or entry point is named. Start by locating the MoE router and router-bias implementation, then trace how expert counts are represented across layers. Done means heterogeneous MoE layers with different numbers of experts are supported without router-bias issues.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.