Verify evals on Papers with Code
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 45/100
- Issue type
- Documentation
- Clarity
- Mostly clear
- Activity status
- Quiet
- Domain
- documentation
Research direction
Start with the linked Papers with Code paper and its listed evaluation and benchmark pages. Verify the imported scores, model names, benchmark protocols, and openness metadata against the paper or official release artifacts; done means reporting any corrections needed for the paper page or repository presentation.
Written by the indexing model from the issue text.
Description
Hi,
Niels here from the open-source team at Hugging Face. Congratulations on your work!
I've made the paper and 7 paper-native evaluations available on Papers with Code.
The paper has results on Coding Agents, Agents, Image Understanding, and 1 additional task page.
The GPT-5.6 Sol results currently rank first on DeepSWE and GPQA Diamond.
The GPT-5.6 Sol (with tools) result currently ranks first on MMMU-Pro.
The GPT-5.6 Sol Ultra result currently ranks first on Terminal Bench 2.1.
Would it be possible to verify these results and let me know if any score, model name, benchmark protocol, or openness metadata should be corrected? The imported rows are tied to the paper or its official release artifacts; comparison-table baselines were not added.
You can also edit the task, methods, project page, and GitHub URL directly from the paper page using your Hugging Face account.
If you'd like to showcase the results in your repository README, you can copy these live leaderboard badges (or use the “Copy PwC badge” button in the Results section):
Kind regards,
Niels
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.5k
- Avg merge
- 1m
- Merged PRs (30d)
- 1k
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from openai/codex
-
enhancement remote
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
bug CLI windows-os
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
macOS sandbox blocks hw.optional.arm64 sysctl, causing Flutter to misdetect Apple Silicon as x64 Openbug CLI sandbox
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
bug CLI TUI
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
-
CLI config enhancement skills
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
kwakseongjae/auto-hwp#319 ·
-
area:cli bug filter-quality good first issue priority:medium
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
bevyengine/bevy#25861 ·
-
comp-datalake
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
ClickHouse/ClickHouse#121222 ·
-
A-linter
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
oxc-project/oxc#26863 ·