Why is intra-document masking only applied for long-context training and not midtraining?
Abierto
- Lenguaje dominante
- Python
- Estrellas
- 1.5k
- Forks
- 315
- Merge medio
- 1 d 9 h
- PR fusionados (30 d)
- 11
Descripción
I noticed that only the long-context stages seem to apply intra-document masking with `generate_doc_lengths=True` in the NumpyDataset config. Is there any reason that midtraining does not also benefit from reducing cross-document attention signals? Thank you!
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.