allenai / allenai/olmes

Reproducing MATH numbers from OLMo-2 release blog

Offen
#11 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
395
Forks
105
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

Thank you for this awesome repo! Seems like a really useful extension of the lm-evaluation-harness. I had a couple of questions regarding the scores from the [OLMo-2 blog](https://allenai.org/blog/olmo2). In particular, under the section "Making OLMo 2 Instruct", the performance of OLMo-2-7B-1124-Instruct on MATH is cited as 31.3.

- Do I understand correctly that this resolves to the minerva_math benchmark?
- Which command can I use to reproduce this? Of all the possible combinations, I got the highest score using the tulu setup, ie. `olmes --model allenai/OLMo-2-1124-7B-Instruct --task "minerva_math::tulu" --output-dir olmes_results/tulu`, shown below:

Summary of primary scores:
minerva_math::tulu: 0.288752
minerva_math_algebra::tulu: 0.506318
minerva_math_counting_and_probability::tulu: 0.291139
minerva_math_geometry::tulu: 0.212944
minerva_math_intermediate_algebra::tulu: 0.125138
minerva_math_number_theory::tulu: 0.227778
minerva_math_prealgebra::tulu: 0.531573
minerva_math_precalculus::tulu: 0.126374

The average of 28.8 is lower than 31.3, so I was wondering if I was doing something wrong here? Furthermore, on a machine with 2x RTXA6000, I get a total processing_time of 108829.94449687004, ie. a bit over 30 hours. Does that seem normal?

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.