NVIDIA / NVIDIA/Megatron-LM

[Feature request] Add Engram memory to HybridModel training

Open
#7,171 0 comments 0 reactions 1 assignee Claimed by @ilml View on GitHub
community-request enhancement waiting-on-maintainers
Dominant language
Python
Stars
17.9k
Forks
4.5k
Avg merge
4d 3h
Merged PRs (30d)
272

Description

**Is your feature request related to a problem? Please describe.**

I independently developed an Engram integration for HybridModel training and would like to
contribute it upstream. It adds n-gram memory before selected standard attention layers.
The main integration challenge is keeping token context and sparse memory state consistent
across pipeline stages, activation recomputation and checkpoints.

**Describe the solution you'd like**

My contribution includes:

- **Layer-owned memory:** each consumer registers its tables, hash state and fusion modules;
hash layouts are determined before pipeline partitioning.
- **Explicit token context:** per-microbatch token inputs reach the relevant pipeline stages
and remain available during recomputation.
- **Distributed memory updates:** local and row-sharded tables, RowSparseAdam, shared
dense/sparse gradient clipping, and native checkpoint save/restore.
- **Training tools:** parallel regression tests and an offline MoE/Engram recipe with MTP1
and TensorBoard/W&B recording. Engram is kept out of the MTP branch.

The initial target is BF16 HybridModel training. Inference, packed sequences, mHC, FP8/FP4,
FSDP, ModelOpt, CUDA graphs, CPU offloading and fine-grained expert overlap are excluded.
The linked PR will provide implementation, tests, usage instructions and the validated
parallel/checkpoint support matrix. I welcome feedback on the interfaces and scope.

**Describe alternatives you've considered**

- **GPTModel:** HybridModel is the upstream direction for new features
([#6071](https://github.com/NVIDIA/Megatron-LM/issues/6071)).
- **Root-owned tables or implicit token caches:** layer ownership and explicit inputs make
pipeline placement and recomputation clearer.
- **Local-only tables:** row sharding also addresses large-table capacity and sparse updates.

[#3382](https://github.com/NVIDIA/Megatron-LM/issues/3382) and
[#3689](https://github.com/NVIDIA/Megatron-LM/pull/3689) are related work only; this integration
does not reuse code from #3689. Its short mock-data tests did not establish the real-data
convergence evidence I needed, so I chose a separate implementation and longer training comparison.

**Additional context**

The algorithm follows [DeepSeek Engram](https://github.com/deepseek-ai/Engram); adapted code
retains its Apache-2.0 attribution.

I ran a FineWeb MoE/Engram comparison using a frozen prototype: **36,754 updates and
9,634,840,576 tokens per model**.

| Main CE | MoE | Engram |
| --- | ---: | ---: |
| Full validation | 2.996060 | 2.971129 |
| Full test | 2.989162 | 2.963933 |

This is one seed with approximately matched backbone parameters; whole-model parameters
and compute are not matched. These prototype results are separate from submission-branch
validation. [Training curves on W&B](https://wandb.ai/yj-tsinghua-university/engram-fineweb)
are publicly readable.

Shared training fixes are reported separately with Engram-free reproductions.

@NVIDIA/mcore-oncall

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.