EleutherAI / EleutherAI/lm-evaluation-harness
How to reproduce the Qwen2.5 base model results on GSM8K Task
Open
- Dominant language
- Python
- Stars
- 14k
- Forks
- 3.6k
- Avg merge
- 4d 2h
- Merged PRs (30d)
- 35
Description
Hi, I am trying to reproduce the Qwen2.5 base model results on GSM8K Task. But I am getting very low scores compared to what was reported in their paper. I noticed that in the GSM8K task files, there is no YAML for reasoning models.
Thanks!
Contributor guide
Assessment
This issue has not been assessed yet.