allenai / allenai/dont-stop-pretraining

Does more steps of pretraining lead to better encoder for downstream tasks?

未关闭
#42 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
544
派生
72
PR 合并指标
30 天内没有已合并 PR

描述

Thank you for your contributions in pretraining. You trained the encoder for 12.5K steps for each domain in pretraining phase before applying the encoder to supervised downstream tasks. Is it possible that the checkpoints that are most suitable for downstream tasks might appear in the middle of the pretraining phase? This phenomenon is obvious in many real applications. Under the circumstances, we might not know if it is a good choice to directly train the model to the maximum step with all of the corpus and take the final checkpoint. Is there any suggestion on that?

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。