agentscope-ai / agentscope-ai/Trinity-RFT

Clipping ratio and response length blow-up on GSM8k with default Trinity GRPO config in just 40 batches

Abierto
#176 10 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
701
Forks
79
Merge medio
8 h 7 min
PR fusionados (30 d)
1

Descripción

Hi!

We're getting response length and clip ratio blow-up on GSM8k which leads to halting of any learning.

We've tried the default Trinity GRPO config `examples/grpo_gsm8k/train_gsm8k.yaml` (except we put the number of epochs to a large value)

By chance, do you have any traces of longer wandb runs published for for the baseline configs GSM8k?

Would you have any stability suggestions on preventing such blow-up / mitigations?

Thank you :)

Image

---

I'm hypothesizing it's some basic parameters like batch size / learning rate / world size / GRPO hyperparams at fault? Or any other learning params?

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.