huggingface / huggingface/alignment-handbook
Add instrutions to evaluate on academic datasets
Open
- Dominant language
- Python
- Stars
- 5.7k
- Forks
- 490
- Avg merge
- 2m
- Merged PRs (30d)
- 1
Description
The paper evaluates on ARC, HellaSwag, MMLU, and TruthfulQA, but this repo does not reference these evals.
Adding short explanation regarding these evals (e.g., in https://github.com/huggingface/alignment-handbook/tree/main/scripts#evaluating-chat-models) would be nice
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.