alibaba / alibaba/ROLL

Example for training on SWE (agentic software-engineering) tasks?

Open
#458 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.4k
Forks
312
Avg merge
1h 2m
Merged PRs (30d)
2

Description

Is there an example/recipe for grpo training on multi-turn SWE-style tasks (SWE-bench / SWE-Gym)?

Looking for:
- Multi-turn agentic rollouts with tool calls (shell, file edits, tests)
- How environment/reward is wired (e.g. tests-passing as reward)
- Setup for long trajectories / large context
- Setup for how docker/podman is handled efficiently during rollout

If nothing SWE-specific exists, a pointer to the closest multi-turn example to adapt would help. Happy to contribute one back.

Thanks!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the closest existing multi-turn example in ROLL and tracing how its tool calls, environment, and reward are wired. Compare it with the requested SWE-bench or SWE-Gym workflow, including long trajectories and Docker or Podman handling. Done means a documented example or recipe covers the requested rollout, tests-passing reward, context, and container setup.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, python
Domain
ai, devtools, testing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.