graykode / graykode/gpt-2-Pytorch
How to train/fine tune the model with multiple GPUs?
Open
- Dominant language
- Python
- Stars
- 1k
- Forks
- 230
- PR merge metrics
- No merged PRs in 30d
Description
I have pulled the code from branch [train](https://github.com/graykode/gpt-2-Pytorch/tree/train). Is there a way to train or fine tune the GPT-2 model with data parallelism on multiple GPUs? Thanks for your help.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.