OpenEuroLLM / OpenEuroLLM/Taskboard
Better LLM judges evaluation for instruction-tuned models
Open
@geoalgo is already working on this.
Since Mar 3, 2026.
- Dominant language
- No language data
- Stars
- 3
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Right now we use mostly Alpaca-Eval, Arena-Hard (english) and m-Arena-Hard for multilingual data.
They suffer for different pitfalls including:
- relying on a fixed baseline (all) which provides scores which are not transitive
- being mostly english (first two), m-arena-hard is multilingual but the quality of translation is very poor
- they provide scores (winrate against a fixed baseline) which are different from ELO ratings obtained on LMArena or ComparIA
We propose to allow to directly estimate ELO ratings and calibrate predictions with conformal predictions to obtain calibrated uncertainty.
We will add this method in openjury.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.