huggingface / huggingface/datablations

Validation loss vs model size per step

Open
#14 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
345
Forks
19
PR merge metrics
No merged PRs in 30d

Description

Hi @Muennighoff
Great paper, very impressive work and very detailed - thanks for releasing the data!
I wonder about a small discrepancy that I see between your work and scaling rules.
I replotted the data in figure 15 for 1 epoch, all 3 models on 1 plot:

![Image](https://github.com/user-attachments/assets/de94cec6-27e0-4896-9a25-7aaeed7bca1b)

You can see in Scaling Rules image that more parameters converge faster and have better loss. But in your experiments it seems that the 9B paraments model behave differently
What are your thoughts about it?
Thanks!

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.