huggingface / huggingface/lighteval

[EVAL] Add Arena Hard

Open
#876 0 comments 0 reactions 0 assignees View on GitHub
science-team
Dominant language
Python
Stars
2.5k
Forks
555
Avg merge
1d 6h
Merged PRs (30d)
1

Description

## Evaluation short description

Similar to AlpacaEval, Arena Hard has been developed by LMSYS to provide an LLM-as-judge benchmark that correlates highly with human preferences (at least those measured on the chatbot arena).

As models have become more capable, this benchmark is becoming a popular proxy for model developer to baseline their algorithms against.

## Evaluation metadata
Provide all available
- Paper url:
- Github url: https://github.com/lmarena/arena-hard-auto
- Dataset url:

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.