anthropics / anthropics/claude-cookbooks

[QUESTION] Request for additional datasets & tool suites to reproduce tool evaluation & optimization gains

Open
#648 0 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Jupyter Notebook
Stars
52.7k
Forks
6.3k
Avg merge
25m
Merged PRs (30d)
6

Description

### Your Question

Hi,
I’ve carefully read the related blog posts and gone through the code in the tool_evaluation cookbook:https://github.com/anthropics/claude-cookbooks/tree/main/tool_evaluation

I’m very interested in Claude’s tool evaluation framework and the practical gains from tool optimization.

To better reproduce and validate the evaluation pipeline and measure improvements from tooling/prompting optimization, I’d like to ask:

Are there publicly available or recommended evaluation datasets for tool use, tool calling, and agentic workflows that work well with this cookbook?Do you provide standard tool suites/simulated tool environments to replicate the evaluation setup shown in the examples?
Any pointers to benchmarks (like τ‑bench, SWE‑bench, or similar) you recommend for end-to-end tool evaluation would also be very helpful.
Thanks for sharing this great cookbook! It’s super useful for building and evaluating tool-using agents.

### Related Notebook (if applicable)

_No response_

### What I've Tried

_No response_

### Additional Context

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.