huggingface / huggingface/cosmopedia

Number of shots during evaluation

Open
#25 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
574
Forks
49
PR merge metrics
No merged PRs in 30d

Description

Thanks for your great work!

From https://github.com/huggingface/cosmopedia/tree/main/evaluation#benchmark-evaluation, is this the exact command you are using for evaluation? Because I found most of them are 0-shot which is inconsistent with the OpenLLM leaderboard.

Could you help confirm this?

Thank you!

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.