huggingface / huggingface/lighteval
[EVAL]: Add more African Benchmarks
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Evaluation short description
- Why is this evaluation interesting?
This focuses on 16 African languages, evaluated on three knowledge QA and reasoning tasks such as AfriMMLU, AfriMGSM and AfriXNLI, human translated from MMLU, MGSM and XNLI respectively.
- How used is it in the community?
## Evaluation metadata: IrokoBench
Provide all available
- Paper url: https://arxiv.org/abs/2406.03368
- Github url:
- Dataset url: https://huggingface.co/collections/masakhane/irokobench-665a21b6d4714ed3f81af3b1
## Evaluation metadata: Uhura
Provide all available
- Paper url:
- Github url:
- Dataset url: https://huggingface.co/datasets/ebayes/uhura-arc-easy-clean , https://huggingface.co/datasets/ebayes/uhura-truthfulqa-clean
## Evaluation metadata: SIB-200
Provide all available
- Paper url: https://aclanthology.org/2024.eacl-long.14/
- Github url: https://github.com/dadelani/sib-200
- Dataset url: https://huggingface.co/datasets/Davlan/sib200
@NathanHB
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.