bigscience-workshop / bigscience-workshop/xmtf

What is the training config?

Open
#6 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
534
Forks
43
PR merge metrics
No merged PRs in 30d

Description

Hello, thanks for your work! I want to try to implement this work myself, but I cann't achieve the high performance by xP3 and mT0-xxl as shown in the paper Crosslingual Generalization through Multitask Finetuning. I wonder the training details of this work, how many steps do you train the model, and what is your lr-decay-ratio? Could I get the config file to implement your result? Thank you very much!

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.