THUDM / THUDM/slime

GRPO training on Qwen3-30b-a3b coder instruct model reward with math error

Open
#1,117 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

When I use smile on GRPO training on Qwen3-30b-a3b coder model , set rm-type to math as below:

ROLLOUT_ARGS=(
--prompt-data $data_path
--input-key prompt
--label-key label
--rollout-shuffle
--rm-type math
--num-rollout 20
--rollout-batch-size 32
--n-samples-per-prompt 8
--rollout-max-response-len 32768
--rollout-temperature 0.8

--global-batch-size 128
--balance-data

(RolloutManager pid=1353, ip=22.3.71.164) [2025-12-15 14:15:52] sglang_rollout.py:367 - First rollout sample: ['[{\'content\': \'Solve the following math problem step by step. The last line of your response should be of the form Answer: \\\\boxed{$Answer} where $Answer is the answer to the problem.\\n\\n方程 x^3-9x^2+8x+2=0 有 3 个实根 p,q,r,则 \\\\df{1}{p^2}+\\\\df{1}{q^2}+\\\\df{1}{r^2}=__________.\\n\\nRemember to put your answer on its own line after "Answer:".\', \'role\': \'user\'}]I need to find the value of $\\frac{1}{p^2} + \\frac{1}{q^2} + \\frac{1}{r^2}$ where $p$, $q$, and $r$ are the roots of $x^3 - 9x^2 + 8x + 2 = 0$.\n\nFirst, I\'ll use Vieta\'s formulas for the polynomial $x^3 - 9x^2 + 8x + 2 = 0$.\n\nFor a cubic polynomial $x^3 + ax^2 + bx + c = 0$ with roots $p$, $q$, $r$:\n- $p + q + r = -a$\n- $pq + pr + qr = b$ \n- $pqr = -c$\n\nFor $x^3 - 9x^2 + 8x + 2 = 0$, we have $a = -9$, $b = 8$, $c = 2$.\n\nTherefore:\n- $p + q + r = 9$\n- $pq + pr + qr = 8$\n- $pqr = -2$\n\nNow I need to find $\\frac{1}{p^2} + \\frac{1}{q^2} + \\frac{1}{r^2}$.\n\nI can rewrite this as:\n$$\\frac{1}{p^2} + \\frac{1}{q^2} + \\frac{1}{r^2} = \\frac{q^2r^2 + p^2r^2 + p^2q^2}{p^2q^2r^2}$$\n\nThe denominator is $(pqr)^2 = (-2)^2 = 4$.\n\nFor the numerator, I need to find $p^2q^2 + p^2r^2 + q^2r^2$.\n\nI can use the identity:\n$(pq + pr + qr)^2 = p^2q^2 + p^2r^2 + q^2r^2 + 2pqr(p + q + r)$\n\nTherefore:\n$p^2q^2 + p^2r^2 + q^2r^2 = (pq + pr + qr)^2 - 2pqr(p + q + r)$\n\nSubstituting the known values:\n$p^2q^2 + p^2r^2 + q^2r^2 = 8^2 - 2(-2)(9) = 64 + 36 = 100$\n\nTherefore:\n$$\\frac{1}{p^2} + \\frac{1}{q^2} + \\frac{1}{r^2} = \\frac{100}{4} = 25$$\n\nAnswer: $\\boxed{25}$<|im_end|>'], label: 25, reward: 0

So, thee reward should be 1 instead of 0.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with sglang_rollout.py around line 367 and trace how the math reward is computed for the supplied GRPO configuration. Reproduce the Qwen3-30b-a3b coder run with the shown problem and label 25; done means the valid boxed answer receives reward 1 instead of 0.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.