agentscope-ai / agentscope-ai/Trinity-RFT
Clipping ratio and response length blow-up on GSM8k with default Trinity GRPO config in just 40 batches
Đang mở
- Ngôn ngữ chính
- Python
- Star
- 701
- Fork
- 79
- Merge trung bình
- 8 giờ 7 phút
- Pull request đã merge (30 ngày)
- 1
Mô tả
Hi!
We're getting response length and clip ratio blow-up on GSM8k which leads to halting of any learning.
We've tried the default Trinity GRPO config `examples/grpo_gsm8k/train_gsm8k.yaml` (except we put the number of epochs to a large value)
By chance, do you have any traces of longer wandb runs published for for the baseline configs GSM8k?
Would you have any stability suggestions on preventing such blow-up / mitigations?
Thank you :)
---
I'm hypothesizing it's some basic parameters like batch size / learning rate / world size / GRPO hyperparams at fault? Or any other learning params?
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.