agentscope-ai / agentscope-ai/Trinity-RFT

Example for training on SWE (agentic software-engineering) tasks?

Open
#573 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
701
Forks
79
Avg merge
8h 7m
Merged PRs (30d)
1

Description

Is there an example/recipe for grpo training on multi-turn SWE-style tasks (SWE-bench / SWE-Gym)?

Looking for:
- Multi-turn agentic rollouts with tool calls (shell, file edits, tests)
- How environment/reward is wired (e.g. tests-passing as reward)
- Setup for long trajectories / large context
- Setup for how docker/podman is handled efficiently during rollout

If nothing SWE-specific exists, a pointer to the closest multi-turn example to adapt would help. Happy to contribute one back.

Thanks!

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.