OpenEuroLLM / OpenEuroLLM/Taskboard
Add pretraining evals for decontamination
Open
@haideraltahan is already working on this.
Since Apr 29, 2026.
T5.1 - static evals
- Dominant language
- No language data
- Stars
- 3
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Eval list: https://docs.google.com/document/d/1lpKgHLgBK8usB6RZmD_8yS0m-eUSuuxJ3DtV7sGeL2g/edit?usp=sharing
- Global MMLU (general knowledge multiple choice)
- INCLUDE base-44 (knowledge- and reasoning multiple choice )
- Flores-200 (sentence translation)
- BeleBele (multiple choice reading comprehension)
- SIB-200 (sentence topic classification) - LM Eval Harness
- Multiblimp (pairwise sentence ranking) - LM Eval harness
- ARC challenge (multiple choice general knowledge) - LM Eval harness
- Hellaswag
- Xcsqa In LighEval
- Global MGSM (math) Google languages are in LightEval, maybe easy to add the others. Omit Greek, as it has a non-commercial license.
- PolyMath (math)
- Global PIQA (multiple choice questions of which many are culturally grounded in local culture - not a parallel benchmark) in LM Eval Harness
- Doclevel MT (doc level translation) report averages for English to x and x to English separately so we can analyse the separately. In LM Eval Harness (I'm using WMT for now, not sure if it's correct
- OpenSubtitles
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.