Example for training on SWE (agentic software-engineering) tasks?
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 312
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 2
Description
Is there an example/recipe for grpo training on multi-turn SWE-style tasks (SWE-bench / SWE-Gym)?
Looking for:
- Multi-turn agentic rollouts with tool calls (shell, file edits, tests)
- How environment/reward is wired (e.g. tests-passing as reward)
- Setup for long trajectories / large context
- Setup for how docker/podman is handled efficiently during rollout
If nothing SWE-specific exists, a pointer to the closest multi-turn example to adapt would help. Happy to contribute one back.
Thanks!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the closest existing multi-turn example in ROLL and tracing how its tool calls, environment, and reward are wired. Compare it with the requested SWE-bench or SWE-Gym workflow, including long trajectories and Docker or Podman handling. Done means a documented example or recipe covers the requested rollout, tests-passing reward, context, and container setup.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- docker, python
- Domain
- ai, devtools, testing
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100