huggingface / huggingface/datablations
A question about the conclusion of this paper
Open
- Dominant language
- Jupyter Notebook
- Stars
- 345
- Forks
- 19
- PR merge metrics
- No merged PRs in 30d
Description
"Scaling Data-Constrained Language Models" is a very nice paper, and I learn a lot from this paper.
However, I have a question about this paper:
In the abstract and Figure 1, it recommends we should train 4 epochs.
But Figure 3 shows that we should choose 59 epochs.
So my question is why the optimal epoch is not 4 epochs in Figure 3.
Thanks in advance.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.