OpenEuroLLM / OpenEuroLLM/Taskboard

Better LLM judges evaluation for instruction-tuned models

Open
#184 0 comments 0 reactions 1 assignee View on GitHub

@geoalgo is already working on this.

Since Mar 3, 2026.

Dominant language
No language data
Stars
3
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Right now we use mostly Alpaca-Eval, Arena-Hard (english) and m-Arena-Hard for multilingual data.

They suffer for different pitfalls including:

  • relying on a fixed baseline (all) which provides scores which are not transitive
  • being mostly english (first two), m-arena-hard is multilingual but the quality of translation is very poor
  • they provide scores (winrate against a fixed baseline) which are different from ELO ratings obtained on LMArena or ComparIA

We propose to allow to directly estimate ELO ratings and calibrate predictions with conformal predictions to obtain calibrated uncertainty.

We will add this method in openjury.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.