huggingface / huggingface/lighteval

[EVAL] Request support for AIME26

Open
#1,167 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
2.5k
Forks
555
Avg merge
1d 6h
Merged PRs (30d)
1

Description

## Evaluation short description

AIME26 is the latest AIME-style math reasoning benchmark, commonly used to evaluate LLM performance on competition-level problems requiring multi-step reasoning and exact answers.

It is widely used in the community as a standard reference benchmark for mathematical reasoning, alongside earlier AIME versions.

## Evaluation metadata

- Paper url: N/A (AIME is an AMC competition benchmark)
- Github url: N/A
- Dataset url: https://huggingface.co/datasets/EleutherAI/aime_2024 (AIME-style datasets; AIME26 variants also available in community repos)

---

Hi LightEval team,

Does LightEval currently support evaluating models on **AIME26**?

If not, is there a recommended way to add it as a custom task, or any plan to support it officially?

Thanks!

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.