NVIDIA / NVIDIA/Megatron-LM

Feature Request: DenseMixer - dense forward pass for MoE router gradient estimation during RL

Open
#4,169 0 comments 0 reactions 0 assignees View on GitHub
enhancement module: rl
Dominant language
Python
Stars
17.9k
Forks
4.5k
Avg merge
4d 3h
Merged PRs (30d)
272

Description

## Summary
Request to investigate supporting [DenseMixer](https://github.com/yaof20/DenseMixer)-style training in Megatron Core's MoE layer. DenseMixer computes all expert outputs during the forward pass and uses a straight-through estimator (STE) to provide more precise gradients to the router, bypassing the non-differentiable Top-K
selection.

The MoE architecture is unchanged — this is a training-mode flag that enables dense expert computation for better router gradient estimation during post-training (SFT/RL).

## Motivation
Standard Top-K routing is non-differentiable — the router only receives gradient signal from selected experts, limiting its ability to learn optimal routing during fine-tuning. DenseMixer addresses this by computing all expert outputs in the forward pass for gradient estimation while preserving sparse Top-K selection at inference.

Results across multiple MoE architectures during SFT:
- Qwen1.5-MoE-A2.7B: +2.2% average across 7 benchmarks
- OLMoE-1B-7B: +2.9% average
- Qwen3-30B-A3B: +3.7% on GPQA-Diamond

The overhead is modest: ~1.46× FLOPs (expert weights are already loaded, so memory overhead is negligible) with 9–29% wall-clock increase depending on dataset size. No inference cost. Compatible with LoRA/PEFT.

## Requested Feature
Investigate adding a configuration flag in Megatron Core's MoE layer to enable dense expert computation with STE-based router gradient estimation during the forward/backward pass. Standard sparse Top-K routing is preserved at inference.

This flag would allow downstream RL frameworks that use Megatron Core as their training backend to opt into improved router training without modifying core MoE internals.

## References
- [DenseMixer (Yao, Cui, Zhang, Liu et al.)](https://github.com/yaof20/DenseMixer)
- [Technical blog](https://fengyao.notion.site/moe-posttraining)

Contributor guide

Open the contributing guide

Research direction

Start by reading Megatron Core's MoE layer and compare its forward and backward routing behavior with the linked DenseMixer reference. Define how a training-mode flag, dense expert computation, and a straight-through router estimator should integrate while preserving sparse Top-K inference; the issue does not name specific files or tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.