huggingface / huggingface/datablations

A question about the conclusion of this paper

Open
#6 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
345
Forks
19
PR merge metrics
No merged PRs in 30d

Description

"Scaling Data-Constrained Language Models" is a very nice paper, and I learn a lot from this paper.

However, I have a question about this paper:

In the abstract and Figure 1, it recommends we should train 4 epochs.

But Figure 3 shows that we should choose 59 epochs.

So my question is why the optimal epoch is not 4 epochs in Figure 3.

Thanks in advance.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.