agentscope-ai / agentscope-ai/Trinity-RFT

Example for training on SWE (agentic software-engineering) tasks?

オープン
#573 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
701
フォーク
79
平均マージ
8時間 7分
マージ済み PR(30日)
1

説明

Is there an example/recipe for grpo training on multi-turn SWE-style tasks (SWE-bench / SWE-Gym)?

Looking for:
- Multi-turn agentic rollouts with tool calls (shell, file edits, tests)
- How environment/reward is wired (e.g. tests-passing as reward)
- Setup for long trajectories / large context
- Setup for how docker/podman is handled efficiently during rollout

If nothing SWE-specific exists, a pointer to the closest multi-turn example to adapt would help. Happy to contribute one back.

Thanks!

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。