DAMO-NLP-SG / DAMO-NLP-SG/TempReason
Question Regarding L3 Training Results
- Dominant language
- Python
- Stars
- 33
- Forks
- 2
- PR merge metrics
- No merged PRs in 30d
Description
Hello! I'm very interested in your work.
I've run the T5-SFT training on the L2 and L3 datasets for ReasonQA and OBQA. However, I've noticed a significant discrepancy between the reproduced L3 results in ReasonQA and the results presented in the paper.
My results show an EM of 28.56 and an F1 score of 42.61, while the paper reports an EM of 78.2 and an F1 score of 83.0.
In my config file, I've set the text to "fact_context," and both the training and test datasets are configured for L3.
Could you please advise me on what I need to modify to align my results with the ones presented in the paper?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.