huggingface / huggingface/lighteval

[EVAL] Add LiveBench

Open
#875 0 comments 0 reactions 0 assignees View on GitHub
science-team
Dominant language
Python
Stars
2.5k
Forks
555
Avg merge
1d 6h
Merged PRs (30d)
1

Description

## Evaluation short description
Most benchmarks become saturated over time as model developers tune their models to those specific capabilities. LiveBench is designed to address this by continually updating the set of prompts used for evaluation and thereby preventing train/test leakage or overfitting. It is also widely used in modern model releases.

## Evaluation metadata
Provide all available
- Paper url: https://livebench.ai/#/
- Github url: https://livebench.ai/#/
- Dataset url: https://huggingface.co/collections/livebench/livebench-67eaef9bb68b45b17a197a98

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.