huggingface / huggingface/alignment-handbook
Question about the evaluation dataset
Open
- Dominant language
- Python
- Stars
- 5.7k
- Forks
- 490
- Avg merge
- 2m
- Merged PRs (30d)
- 1
Description
Hi, I wonder which TruthfulQA task you are focusing on during evaluation? MC1, MC2, or generation task?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.