huggingface / huggingface/lighteval

[EVAL] Add ArmBench-LLM Armenian evaluation suite

Open
#1,306 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
2.5k
Forks
555
Avg merge
1d 6h
Merged PRs (30d)
1

Description

## Evaluation short description

* **Why is this evaluation interesting?**
ArmBench-LLM adds standardized Armenian-language evaluation across 24 tasks, including classification, QA, reasoning, summarization, translation, NER, and language correction. Armenian currently has very limited coverage in common LLM evaluation frameworks.

* **How used is it in the community?**
The benchmark was released recently, so adoption is still growing. Its dataset, results, and evaluation code are publicly available, and the proposed LightEval integration is implemented in PR #1289.

## Evaluation metadata

* Paper/report URL: https://huggingface.co/blog/Metric-AI/armbench-llm
* GitHub/implementation URL: https://github.com/huggingface/lighteval/pull/1289
* Dataset URL: https://huggingface.co/datasets/Metric-AI/ArmBench-LLM-data

Related PR: #1289

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading PR #1289, the ArmBench-LLM report, and the linked dataset to understand the proposed LightEval integration and its 24 tasks. Compare the implementation with the issue metadata and confirm that the Armenian evaluation suite is fully represented and usable in the framework.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface, python
Domain
machine-learning, testing-qa
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.