huggingface / huggingface/open-r1

Bugs in evaluation? Cannot reproduce the numbers for deepseek models

Open
#97 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
26.5k
Forks
2.5k
PR merge metrics
No merged PRs in 30d

Description

For DeepSeek-R1-Distill-Qwen-1.5B, I got 30.0 for AIME, which is good (28.9 in the paper)

For DeepSeek-R1-Distill-Qwen-7B, I got 50.0 for AIME, which is 5 points lower than 55.5 in the paper.

For DeepSeek-R1-Distill-Llama-8B, I only got 23 for AIME, but it is 50.4 in the paper

I don't change anything. All hyperparameters are default. Is this a bug or something?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the evaluation entry point and the default configuration used for AIME, then reproduce the reported results for the three DeepSeek-R1-Distill models. Compare the evaluation setup with the paper's setup; done means identifying the source of the discrepancy or documenting why the numbers differ.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.