NVIDIA-NeMo / NVIDIA-NeMo/Gym

Produce evaluation results for NOOA agents with Nemotron 3.5

Open
#3,457 0 comments 0 reactions 1 assignee View on GitHub

@init-nikhil is already working on this.

Since Sep 17, 2026.

area:agent area:evaluation dependency-CLU feature
Dominant language
Python
Stars
1.2k
Forks
349
Avg merge
1d 23h
Merged PRs (30d)
324

Description

Parent

Sub-issue of #3453.

Goal

Produce reproducible Nemotron 3.5 evaluation results for NOOA agents on DeepSWE and SWE Bench Pro.

Scope

  • Run the DeepSWE and SWE Bench Pro evaluations with the sandboxed NOOA runner and Nemotron 3.5.
  • Preserve model, agent, benchmark, and sampling configuration.
  • Preserve rollout artifacts, rewards, completion status, and failure details.
  • Summarize results and identify follow-up work.

Acceptance criteria

  • DeepSWE evaluation results are recorded.
  • SWE Bench Pro evaluation results are recorded.
  • Model, agent, benchmark, and sampling configuration are reproducible.
  • Results include rewards, completion rates, and categorized failures.
  • Follow-up implementation gaps are linked or filed separately.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.