huggingface / huggingface/lighteval
[EVAL] Add Arena Hard
Open
science-team
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Evaluation short description
Similar to AlpacaEval, Arena Hard has been developed by LMSYS to provide an LLM-as-judge benchmark that correlates highly with human preferences (at least those measured on the chatbot arena).
As models have become more capable, this benchmark is becoming a popular proxy for model developer to baseline their algorithms against.
## Evaluation metadata
Provide all available
- Paper url:
- Github url: https://github.com/lmarena/arena-hard-auto
- Dataset url:
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.