google-deepmind / google-deepmind/superhuman
Difficulty levels for IMO-AnswerBench still not shared
- Dominant language
- Lean
- Stars
- 800
- Forks
- 85
- PR merge metrics
- No merged PRs in 30d
Description
Dear Authors,
Thank you for your great work on IMO-Bench and the paper 'Towards Robust Mathematical Reasoning'. We would like to replicate the results on IMO-AnswerBench stratified by difficulty, but the difficulty levels, or a mapping to derive them, are still not shared with the community several months after the preprint came out. Could you please update us on the progress on this, if any internal review is currently pending?
I am aware there is a previous Github issue asking the same thing, dated January 2026. However, no response from the authors appears to have been given yet.
Many thanks in advance, and we look forward to using your great benchmark!
@dawsenhwang @nurijunsu
Contributor guide
Research direction
The issue concerns IMO-AnswerBench and requests publicly shared difficulty levels or a mapping to derive them, but it names no repository file, test, or entry point. Start by locating the benchmark data and related published materials, then check whether difficulty metadata or a derivation mapping is available. Done means stratified results can be reproduced from the shared information.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100