huggingface / huggingface/trl

[Feature Request] support dynamic sampling for GRPO trainer

Open
#3,708 4 comments 5 reactions 1 assignee Claimed by @AmineDiro View on GitHub
✨ enhancement 🏋 GRPO
Dominant language
Python
Stars
19.3k
Forks
3k
Avg merge
1d 20h
Merged PRs (30d)
194

Description

### Feature request

Filter most easy and hard samples can increased the diversity of examples, which is beneficial for fine-tuning performance.

trl can provide a custom function between generation and computing per_token_logps for user to evaluate the quality of prompt, to determinate whether current sample be used in next step or not.

### Motivation

- dynamic sampling from [DAPO](https://arxiv.org/abs/2503.14476), page 5, 3.2 section.
- Rollout Rescue Mechanism and Intra-Batch Informative Substitution in [Polaris](https://hkunlp.github.io/blog/2025/Polaris/)

### Your contribution

I'm pleased to contribute this feature

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.