Why is intra-document masking only applied for long-context training and not midtraining?
未關閉
- 主要語言
- Python
- 星號
- 1.5k
- 分支
- 315
- 平均合併
- 1 天 9 小時
- 30 天內合併 PR
- 11
描述
I noticed that only the long-context stages seem to apply intra-document masking with `generate_doc_lengths=True` in the NumpyDataset config. Is there any reason that midtraining does not also benefit from reducing cross-document attention signals? Thank you!
貢獻指南
評估
這個 Issue 還沒有評估資料。