google-deepmind / google-deepmind/mathematics_dataset
Questions about the training settings
- Dominant language
- Python
- Stars
- 2k
- Forks
- 275
- PR merge metrics
- No merged PRs in 30d
Description
Hi! I am really interested in this fascinating work. However, I have some questions about the training methods for the transformer model.
In the paper you mention the transformer model is trained with learning rate = 6e-4 but do not say which lr decay method you are using, which I am curious about. I am also curious about the number of layers in the encoder and decoder.
Could you please demonstrate more specifically about the training settings? It will be more convenient for someone like me who want to reproduce your results if you could just publish your training source codes.
Thank you very much!
Contributor guide
Assessment
This issue has not been assessed yet.