huggingface / huggingface/lighteval
[FT] Load entire benchmark (data + spec) from the hub
Open
feature
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Issue encountered
Having the ability for benchmark builders to create a dataset and a spec that will be read with lighteval.
## Solution/Feature
A dataset, with the benchmark data
A python file that defines the prompt metric etc OR a yaml file if you want to use regular metrics.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.