bethgelab / bethgelab/sober-reasoning

The code (command) to reproduce the official benchmark

Open
#18 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
HTML
Stars
92
Forks
6
PR merge metrics
No merged PRs in 30d

Description

Hi,

Thanks for the nice framework.

What is the exact command to evaluate a new model to compare it fairly with the models shown on the official leaderboard?

Contributor guide

No contributing guide indexed for this repository

Research direction

No file or test is named in the issue. First inspect the repository's benchmark and leaderboard instructions and identify the command used for official evaluations. Document the reproducible command and required inputs so a new model can be compared fairly with the leaderboard models.

Written by the indexing model from the issue text.

Assessment

Domain
documentation, machine-learning
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.