huggingface / huggingface/evaluation-guidebook
New Alternative to LLM-as-a-judge!
- Dominant language
- Jupyter Notebook
- Stars
- 2.1k
- Forks
- 125
- PR merge metrics
- No merged PRs in 30d
Description
Hello Clementine and the Evaluation Community,
We would like to introduce you to our **new** metric, **HumanRankEval**, an alternative to the popular 'llm-as-a-judge'. Instead of using the LLM to judge machine-generated text, we use human-generated text to 'judge' the LLM! :) Please take a look, thank you very much! Let us know what you think :)
NAACL '24 PAPER LINK: https://aclanthology.org/2024.naacl-long.456/
CODE: https://github.com/huawei-noah/noah-research/tree/master/NLP/HumanRankEval
DATA: https://huggingface.co/datasets/huawei-noah/human_rank_eval
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.