huggingface / huggingface/open-r1
Bugs in evaluation? Cannot reproduce the numbers for deepseek models
- Dominant language
- Python
- Stars
- 26.5k
- Forks
- 2.5k
- PR merge metrics
- No merged PRs in 30d
Description
For DeepSeek-R1-Distill-Qwen-1.5B, I got 30.0 for AIME, which is good (28.9 in the paper)
For DeepSeek-R1-Distill-Qwen-7B, I got 50.0 for AIME, which is 5 points lower than 55.5 in the paper.
For DeepSeek-R1-Distill-Llama-8B, I only got 23 for AIME, but it is 50.4 in the paper
I don't change anything. All hyperparameters are default. Is this a bug or something?
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the evaluation entry point and the default configuration used for AIME, then reproduce the reported results for the three DeepSeek-R1-Distill models. Compare the evaluation setup with the paper's setup; done means identifying the source of the discrepancy or documenting why the numbers differ.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100