huggingface / huggingface/lighteval

[EVAL] Add BFCL-v3

Open
#873 0 comments 1 reaction 0 assignees View on GitHub
science-team
Dominant language
Python
Stars
2.5k
Forks
555
Avg merge
1d 6h
Merged PRs (30d)
1

Description

## Evaluation short description

2025 is apparently the "year of agents" and as a result agentic evals are becoming the next frontier to measure model capabilities. BFCL-v3 is the most popular benchmark to evaluate the **tool calling** capabilities of LLMs and is now reported in most model releases.

## Evaluation metadata
Provide all available
- Paper url: https://gorilla.cs.berkeley.edu/leaderboard.html
- Github url: https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard
- Dataset url: https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard/bfcl_eval/data

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.