huggingface / huggingface/lighteval
[EVAL] Big-Bench Extra Hard (BBEH)
Open
good first issue
help wanted
new-task
science-team
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Evaluation short description
Google has releases BBEH as a way to compensate for the saturation of BBH in the latest generation of LLMs. Overall looks like a good benchmark to probe reasoning capabilities.
## Evaluation metadata
Provide all available
- Paper url: https://arxiv.org/pdf/2502.19187
- Github url: https://github.com/google-deepmind/bbeh
- Dataset url:
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.