bethgelab / bethgelab/sober-reasoning
The code (command) to reproduce the official benchmark
Open
- Dominant language
- HTML
- Stars
- 92
- Forks
- 6
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
Thanks for the nice framework.
What is the exact command to evaluate a new model to compare it fairly with the models shown on the official leaderboard?
Contributor guide
No contributing guide indexed for this repository
Research direction
No file or test is named in the issue. First inspect the repository's benchmark and leaderboard instructions and identify the command used for official evaluations. Document the reproducible command and required inputs so a new model can be compared fairly with the leaderboard models.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation, machine-learning
- Issue type
- Documentation
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100