huggingface / huggingface/open-r1

The score in math_500 is much lower than the model released by deepseek.

Open
#369 9 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
26.5k
Forks
2.5k
PR merge metrics
No merged PRs in 30d

Description

I trained Qwen2.5-1.5B on the dataset of OpenR1-Math-220k, but I get the score of 0.522 on math_500. Is this phenomenon expected?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.