huggingface / huggingface/lighteval
[EVAL] Add LiveBench
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Evaluation short description
Most benchmarks become saturated over time as model developers tune their models to those specific capabilities. LiveBench is designed to address this by continually updating the set of prompts used for evaluation and thereby preventing train/test leakage or overfitting. It is also widely used in modern model releases.
## Evaluation metadata
Provide all available
- Paper url: https://livebench.ai/#/
- Github url: https://livebench.ai/#/
- Dataset url: https://huggingface.co/collections/livebench/livebench-67eaef9bb68b45b17a197a98
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.