facebookresearch / facebookresearch/fairseq2
Fairseq2 AdamW GPU memory consumption
Open
bug
- Dominant language
- Python
- Stars
- 1.1k
- Forks
- 144
- Avg merge
- 4d 1h
- Merged PRs (30d)
- 1
Description
Attempts to switch M4T model training/finetuning recipes to fairseq2 AdamW work only with a significant reduction of batch size
Contributor guide
Assessment
This issue has not been assessed yet.