agentscope-ai / agentscope-ai/Trinity-RFT
Example for training on SWE (agentic software-engineering) tasks?
Đang mở
- Ngôn ngữ chính
- Python
- Star
- 701
- Fork
- 79
- Merge trung bình
- 8 giờ 7 phút
- Pull request đã merge (30 ngày)
- 1
Mô tả
Is there an example/recipe for grpo training on multi-turn SWE-style tasks (SWE-bench / SWE-Gym)?
Looking for:
- Multi-turn agentic rollouts with tool calls (shell, file edits, tests)
- How environment/reward is wired (e.g. tests-passing as reward)
- Setup for long trajectories / large context
- Setup for how docker/podman is handled efficiently during rollout
If nothing SWE-specific exists, a pointer to the closest multi-turn example to adapt would help. Happy to contribute one back.
Thanks!
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.