google / google/tunix

Support configurable advantage estimation in AgenticRLLearner (RLOO, DrGRPO)

Open
#1,378 3 comments 0 reactions 1 assignee Claimed by @sizhit2 View on GitHub
type:feature/enhancement
Dominant language
Python
Stars
2.5k
Forks
345
Avg merge
1d 7h
Merged PRs (30d)
240

Description

## Problem

The agentic RL learner (`tunix/rl/agentic/agentic_grpo_learner.py`) currently hardcodes GRPO-style advantage computation. It does not use the `advantage_estimator` field from `AlgorithmConfig` or the function registry, meaning alternative estimators like RLOO and DrGRPO cannot be used for multi turn agentic training (tool-use, reasoning chains, etc.).

This is a significant limitation because:
- RLOO's lower variance baseline is especially valuable in agentic settings where trajectory rewards are noisy due to tool call stochasticity
- DrGRPO's unnormalized advantages can be beneficial when reward distributions shift across agentic episodes
- The non agentic GRPO learner already supports pluggable advantage estimators via `function_registry.get_advantage_estimator()`, but the agentic learner bypasses this

## Proposed Solution

Refactor `AgenticGRPOLearner._compute_advantages()` to route through `function_registry.get_advantage_estimator(self.algo_config.advantage_estimator)` instead of computing group relative advantages inline. This would:

1. Enable RLOO, DrGRPO, and any future estimators for agentic RL with zero additional code
2. Align the agentic and non agentic learner codepaths
3. Preserve backward compatibility (default remains `"grpo"`)

## References

- Non agentic advantage routing: `tunix/rl/grpo/grpo_learner.py` lines 307-312
- Agentic hardcoded computation: `tunix/rl/agentic/agentic_grpo_learner.py`
- RLOO learner PR: #1377
- RLOO paper: https://arxiv.org/abs/2402.14740

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.