huggingface / huggingface/lighteval
[EVAL] Add tau-bench
Open
science-team
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 555
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 1
Description
## Evaluation short description
Similar to BFCL in #873, tau-bench is a popular agentic benchmark that is used to measure the ability of LLMs to use tools in real world domains (e.g. booking flights). It is very popular and reported in modern releases, including those from frontier labs.
We could focus on the τ^2 variant that is more recent.
## Evaluation metadata
Provide all available
- Paper url: https://arxiv.org/abs/2406.12045
- Github url: https://github.com/sierra-research/tau-bench
- Dataset url: simulated
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.