huggingface / huggingface/lighteval
[EVAL] Add ArmBench-LLM Armenian evaluation suite
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Evaluation short description
* **Why is this evaluation interesting?**
ArmBench-LLM adds standardized Armenian-language evaluation across 24 tasks, including classification, QA, reasoning, summarization, translation, NER, and language correction. Armenian currently has very limited coverage in common LLM evaluation frameworks.
* **How used is it in the community?**
The benchmark was released recently, so adoption is still growing. Its dataset, results, and evaluation code are publicly available, and the proposed LightEval integration is implemented in PR #1289.
## Evaluation metadata
* Paper/report URL: https://huggingface.co/blog/Metric-AI/armbench-llm
* GitHub/implementation URL: https://github.com/huggingface/lighteval/pull/1289
* Dataset URL: https://huggingface.co/datasets/Metric-AI/ArmBench-LLM-data
Related PR: #1289
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading PR #1289, the ArmBench-LLM report, and the linked dataset to understand the proposed LightEval integration and its 24 tasks. Compare the implementation with the issue metadata and confirm that the Armenian evaluation suite is fully represented and usable in the framework.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface, python
- Domain
- machine-learning, testing-qa
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100