anthropics / anthropics/claude-cookbooks
[QUESTION] Request for additional datasets & tool suites to reproduce tool evaluation & optimization gains
- Dominant language
- Jupyter Notebook
- Stars
- 52.7k
- Forks
- 6.3k
- Avg merge
- 25m
- Merged PRs (30d)
- 6
Description
### Your Question
Hi,
I’ve carefully read the related blog posts and gone through the code in the tool_evaluation cookbook:https://github.com/anthropics/claude-cookbooks/tree/main/tool_evaluation
I’m very interested in Claude’s tool evaluation framework and the practical gains from tool optimization.
To better reproduce and validate the evaluation pipeline and measure improvements from tooling/prompting optimization, I’d like to ask:
Are there publicly available or recommended evaluation datasets for tool use, tool calling, and agentic workflows that work well with this cookbook?Do you provide standard tool suites/simulated tool environments to replicate the evaluation setup shown in the examples?
Any pointers to benchmarks (like τ‑bench, SWE‑bench, or similar) you recommend for end-to-end tool evaluation would also be very helpful.
Thanks for sharing this great cookbook! It’s super useful for building and evaluating tool-using agents.
### Related Notebook (if applicable)
_No response_
### What I've Tried
_No response_
### Additional Context
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.