google-deepmind / google-deepmind/gemma

Reproducing evaluations

Open
#42 2 comments 0 reactions 1 assignee Claimed by @tilakrayal View on GitHub
question
Dominant language
Python
Stars
5.7k
Forks
1k
Avg merge
10h 33m
Merged PRs (30d)
2

Description

Trying to reproduce evaluation numbers but not able to.

Ex : For gemma-2-9b, the technical report mentions 68.2 on BBH 3 shot CoT while the open llm [leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard) reported 5.05

Was there any special setting used for evaluations ?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.