allenai / allenai/dont-stop-pretraining
Does more steps of pretraining lead to better encoder for downstream tasks?
未关闭
- 主要语言
- Python
- 星标
- 544
- 派生
- 72
- PR 合并指标
- 30 天内没有已合并 PR
描述
Thank you for your contributions in pretraining. You trained the encoder for 12.5K steps for each domain in pretraining phase before applying the encoder to supervised downstream tasks. Is it possible that the checkpoints that are most suitable for downstream tasks might appear in the middle of the pretraining phase? This phenomenon is obvious in many real applications. Under the circumstances, we might not know if it is a good choice to directly train the model to the maximum step with all of the corpus and take the final checkpoint. Is there any suggestion on that?
贡献指南
这个仓库没有索引到贡献指南
评估
这个 Issue 还没有评估数据。