Produce evaluation results for NOOA agents with Nemotron 3.5
Open
@init-nikhil is already working on this.
Since Sep 17, 2026.
area:agent
area:evaluation
dependency-CLU
feature
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 349
- Avg merge
- 1d 23h
- Merged PRs (30d)
- 324
Description
Parent
Sub-issue of #3453.
Goal
Produce reproducible Nemotron 3.5 evaluation results for NOOA agents on DeepSWE and SWE Bench Pro.
Scope
- Run the DeepSWE and SWE Bench Pro evaluations with the sandboxed NOOA runner and Nemotron 3.5.
- Preserve model, agent, benchmark, and sampling configuration.
- Preserve rollout artifacts, rewards, completion status, and failure details.
- Summarize results and identify follow-up work.
Acceptance criteria
- DeepSWE evaluation results are recorded.
- SWE Bench Pro evaluation results are recorded.
- Model, agent, benchmark, and sampling configuration are reproducible.
- Results include rewards, completion rates, and categorized failures.
- Follow-up implementation gaps are linked or filed separately.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.