ToolUsage retries a failed tool call twice per outer attempt, not once (schema-args-filter fallback in _use/_ause)
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 58.8k
- Forks
- 8.5k
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 109
Description
Summary
ToolUsage._use/_ause (crewai/tools/tool_usage.py) has two independent retry paths, not one. The outer loop (_run_attempts / _max_parsing_attempts, default 3) is the retry surface I expected. But inside a single outer attempt, there's a second, inner try/except around the actual tool call:
if calling.arguments:
try:
acceptable_args = tool.args_schema.model_json_schema()["properties"].keys()
arguments = {k: v for k, v in calling.arguments.items() if k in acceptable_args}
result = tool.invoke(input=arguments, config=fingerprint_config)
except Exception:
arguments = calling.arguments
result = tool.invoke(input=arguments, config=fingerprint_config)
(sync path, lines 613-629; the async path in _ause is the same shape, lines 356-372.)
This except Exception catches any exception from the first tool.invoke() call — not just a schema/argument-filtering problem. If the tool's own function body raises (a real runtime error, not an argument mismatch), this block re-invokes tool.invoke() a second time with unfiltered arguments, inside the same outer attempt. So every outer attempt calls the tool's underlying function twice, not once — and with the default _max_parsing_attempts=3, a single tool call that always fails ends up invoking the tool 6 times total, not 3.
For a tool with an idempotent effect this is harmless (just wasted calls). For a tool with a non-idempotent side effect (writes a row, sends a request, increments a counter) that happens to raise after the effect lands, this doubles the duplicate-effect exposure of the outer retry loop.
Repro (crewai==1.15.21, no LLM needed — this reaches ToolUsage directly, downstream of tool-call parsing):
import os
os.environ.setdefault("OTEL_SDK_DISABLED", "true")
os.environ.setdefault("CREWAI_DISABLE_TELEMETRY", "true")
from crewai.tools.structured_tool import CrewStructuredTool
from crewai.tools.tool_calling import ToolCalling
from crewai.tools.tool_usage import ToolUsage
calls = []
def my_tool(logical_id: str) -> str:
calls.append(logical_id)
raise RuntimeError("boom")
tool = CrewStructuredTool.from_function(func=my_tool, name="my_tool", description="demo tool")
class _Action:
tool = "my_tool"
tool_input = {"logical_id": "abc"}
usage = ToolUsage(
tools_handler=None, tools=[tool], task=None,
function_calling_llm=None, agent=None, action=_Action(),
)
usage._max_parsing_attempts = 3
calling = ToolCalling(tool_name="my_tool", arguments={"logical_id": "abc"})
usage.use(calling=calling, tool_string="my_tool(logical_id=abc)")
print(f"python-level invocations: {len(calls)}") # 6
print(f"outer run_attempts reported: {usage._run_attempts}") # 4
Output (independently re-run against a fresh install before filing this):
python-level invocations: 6
outer run_attempts reported: 4
Expected vs actual
I'd expect _max_parsing_attempts=3 to mean the tool function is called at most 3 times for a call that keeps failing (matching what _run_attempts reports). Instead it's called up to 6 times, because the inner fallback fires independently on any exception, not only ones that indicate the first arguments dict was wrong.
Context
Found this while building a standalone worked example reproducing a duplicate-effect scenario for a tool with a real side effect (SQLite write) that fails after the effect commits — the inner fallback showed up as an unexpected 2x multiplier on the duplicate count. Repro and full writeup: https://github.com/giskard09/argentum-core/tree/main/examples/conformance/crewai-unguarded-retry (not required to reproduce this — the snippet above is self-contained).
Not sure if the inner fallback's intent was "retry only on schema-filtering exceptions" and it's accidentally too broad, or if it's deliberate — happy to be pointed at the right context if this is expected behavior.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read crewai/tools/tool_usage.py at _use and _ause, focusing on the inner exception fallback and the outer attempt counters. Run the self-contained reproduction from the issue, then add or update regression coverage so a consistently failing tool is invoked once per outer attempt and the reported attempt count remains consistent.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 74/100