huggingface / huggingface/datablations

wonder if LR=1e-3 for mup is optimal value from small-scale proxy model and dropout is crucial for multi-epoch

Open
#13 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
345
Forks
19
PR merge metrics
No merged PRs in 30d

Description

hi authors, thanks for the great work!
i just wonder if LR=1e-3 for mup is optimal value from small-scale proxy model
and how dropout is critical for multi-epoch training.
for the latter, i guess you guys set dropout as 0.1 for regularization but there is no dropout ablation study.
because it's common to set dropout as 0.0 in modern LLM, it would be interesting to know when dropout becomes important

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.