EleutherAI / EleutherAI/lm-evaluation-harness

How to reproduce the Qwen2.5 base model results on GSM8K Task

Open
#3,003 1 comment 5 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
14k
Forks
3.6k
Avg merge
4d 2h
Merged PRs (30d)
35

Description

Hi, I am trying to reproduce the Qwen2.5 base model results on GSM8K Task. But I am getting very low scores compared to what was reported in their paper. I noticed that in the GSM8K task files, there is no YAML for reasoning models.
Thanks!

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.