agentscope-ai / agentscope-ai/Trinity-RFT

Clipping ratio and response length blow-up on GSM8k with default Trinity GRPO config in just 40 batches

オープン
#176 コメント 10 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
701
フォーク
79
平均マージ
8時間 7分
マージ済み PR(30日)
1

説明

Hi!

We're getting response length and clip ratio blow-up on GSM8k which leads to halting of any learning.

We've tried the default Trinity GRPO config `examples/grpo_gsm8k/train_gsm8k.yaml` (except we put the number of epochs to a large value)

By chance, do you have any traces of longer wandb runs published for for the baseline configs GSM8k?

Would you have any stability suggestions on preventing such blow-up / mitigations?

Thank you :)

Image

---

I'm hypothesizing it's some basic parameters like batch size / learning rate / world size / GRPO hyperparams at fault? Or any other learning params?

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。