huggingface / huggingface/lighteval
[EVAL] Add BFCL-v3
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Evaluation short description
2025 is apparently the "year of agents" and as a result agentic evals are becoming the next frontier to measure model capabilities. BFCL-v3 is the most popular benchmark to evaluate the **tool calling** capabilities of LLMs and is now reported in most model releases.
## Evaluation metadata
Provide all available
- Paper url: https://gorilla.cs.berkeley.edu/leaderboard.html
- Github url: https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard
- Dataset url: https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard/bfcl_eval/data
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.