bigscience-workshop / bigscience-workshop/xmtf
What is the training config?
Open
- Dominant language
- Jupyter Notebook
- Stars
- 534
- Forks
- 43
- PR merge metrics
- No merged PRs in 30d
Description
Hello, thanks for your work! I want to try to implement this work myself, but I cann't achieve the high performance by xP3 and mT0-xxl as shown in the paper Crosslingual Generalization through Multitask Finetuning. I wonder the training details of this work, how many steps do you train the model, and what is your lr-decay-ratio? Could I get the config file to implement your result? Thank you very much!
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.