NVIDIA-NeMo / NVIDIA-NeMo/Emerging-Optimizers
Add support for the normalized and orthogonal optimizers with matrix sign
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 274
- Forks
- 51
- Avg merge
- 1d 3h
- Merged PRs (30d)
- 6
Description
Is your feature request related to a problem? Please describe.
The optimizer for normalized layers is introduced by Franz Cesista in https://leloykun.github.io/ponder/steepest-descent-stiefel/#6-bonus-a-muon-like-optimizer-for-the-embedding-and-unembedding-layers
and the optimizer for stiefel is introduced by Jianlin Su in https://kexue.fm/archives/11221
Further, Tilde has introduced a set of optimizers relaxing the strong constraints: https://www.tilderesearch.com/vignettes/gram-space
Describe the solution you'd like
A clear and concise description of what you want to happen.
Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.
Additional context
Add any other context or screenshots about the feature request here.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading the linked Muon-like, Stiefel, and Gram-space optimizer references to determine the intended normalized and orthogonal behaviors. Then inspect the repository's existing optimizer entry points and tests, which are not named in the issue; done means the requested matrix-sign-based optimizers are implemented with coverage for their expected behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100