microsoft / microsoft/foundry-local

Guide + patch: running an OpenAI-compatible tool-calling agent (Hermes) on Foundry Local, and the bugs blocking it (tool_choice=auto, #874)

Open
#875 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
2.6k
Forks
369
Avg merge
2d 17h
Merged PRs (30d)
39

Description

Running an OpenAI-compatible agent (Hermes) on Foundry Local — end-to-end guide, blocking bugs, and a working client-side patch

This is a detailed, reproducible write-up of wiring a real tool-calling agent — Hermes — to Foundry Local as its LLM backend, doing agentic work (file/terminal tool calls). It documents the exact setup, the FL bugs that block it out of the box, and the client-side patch that makes it work today. The goal is to get these gaps fixed in FL so agents work against /v1/chat/completions with no client-side workaround.

Related issues/PRs already filed:

  • #874 — tool_choice="auto" returns tool calls as text with empty tool_calls (the core blocker for agents)
  • #872 (fixes #864) — SDK: opaque pydantic error on malformed tool call
  • #873 — GPU (*-cuda-gpu) variants produce corrupted output on H200/Hopper (forces CPU models)
Environment
  • Foundry Local 0.10.0 (also reproduced on 0.10.1), Linux (Azure Linux 3.0), OpenAI-compatible endpoint on http://127.0.0.1:39839/v1
  • Agent: Hermes CLI (Python), configured with provider: custom pointing at FL
  • Model: qwen2.5-coder-7b-instruct-generic-cpu (a CPU model — see Bug 3)

Part 1 — Setup

1a. Start Foundry Local with a pinned port
foundry server start --port 39839 --idle-timeout 0

--idle-timeout 0 keeps models resident so the agent doesn't pay a reload on every turn. A pinned port is required because the config below hard-codes the base URL. (Note: foundry config set port <N> throws an internal error in 0.10.0, so the port must be pinned on the server start command instead — a separate minor bug.)

1b. Load a tool-capable CPU model
foundry model load qwen2.5-coder-7b-instruct-generic-cpu
1c. Point the agent at FL (Hermes ~/.hermes/config.yaml)
model:
  default: qwen2.5-coder-7b-instruct-generic-cpu
  provider: custom
  base_url: http://127.0.0.1:39839/v1
  max_tokens: 4096          # REQUIRED — see Bug 4
1d. Run the agent
# plain chat
hermes -z 'What is 17 times 3?'

# agentic tool call (file / terminal toolsets)
hermes --yolo -t file     -z 'Create /tmp/test.txt containing hello'
hermes --yolo -t terminal -z 'List the files in the current directory'

With the patch in Part 3 applied, this works end-to-end: the agent plans, FL emits the tool call, the agent executes it, and returns the result.


Part 2 — Bugs that block agentic use out of the box

Bug 1 (blocker) — tool_choice="auto" returns tool calls as text, not structured tool_calls → #874

Agents send tool_choice="auto" (the OpenAI default) so the model can either answer or call a tool. Under auto, FL returns the tool call as text inside message.content with an empty tool_calls array and finish_reason="stop". Only tool_choice="required" yields structured tool_calls.

Reproduction (only tool_choice differs), model qwen2.5-coder-7b-instruct-generic-cpu, temperature=0:

curl -s http://127.0.0.1:39839/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "qwen2.5-coder-7b-instruct-generic-cpu",
  "messages": [{"role":"user","content":"What is the weather in Paris? Use the tool."}],
  "tools": [{"type":"function","function":{"name":"get_weather","description":"Get weather for a city","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}],
  "tool_choice": "auto",
  "temperature": 0
}'

tool_choice="auto" (broken for agents):

finish_reason: stop
tool_calls:    []
content:       "```xml\n<tools>\n    {\"name\": \"get_weather\", \"arguments\": {\"city\": \"Paris\"}}\n</tools>\n```"

tool_choice="required" (structured, as expected):

finish_reason: tool_calls
tool_calls:    [{"index":0,"id":"09xikmhSQ","type":"function","function":{"name":"get_weather","arguments":"{\n  \"city\": \"Paris\"\n}"}}]

Impact: any standard OpenAI agent loop reads message.tool_calls, sees it empty, and treats the tool call as a final text answer — so the agent stalls instead of executing the tool. Forcing tool_choice="required" is not a fix, because it prevents the model from answering directly when no tool is needed.

Requested fix: under auto, run the same server-side tool-call extraction that required already does and populate message.tool_calls with finish_reason="tool_calls".

Bug 2 (related, SDK) — opaque pydantic error on malformed tool call → #864 / PR #872

When a small model emits a malformed tool call (e.g. missing function.name), the Python SDK's complete_chat raises an opaque pydantic ValidationError instead of a FoundryLocalException. Fixed in PR #872 by dropping only the malformed entries and wrapping validation.

Bug 3 — GPU model variants corrupt output on H200/Hopper → #873

Every *-cuda-gpu variant produces deterministic token-doubling/gibberish on H200 (CPU variants are clean), so agents must use CPU models today (slower). Still open; not fixed by 0.10.1.

Bug 4 (client-side, Hermes) — custom-provider default max_tokens overflows FL's context

Hermes' custom provider defaults max_tokens to 65536. FL rejects requests whose max_tokens exceeds the model context (e.g. 32768), surfacing as an incomplete chunked read. Setting model.max_tokens: 4096 (Part 1c) avoids it. Included here so others don't hit it during setup.


Part 3 — The client-side patch that makes it work today

Because Bug 1 lives in FL's closed server, the practical workaround is a client-side adapter: when the provider is the local FL endpoint and tool_calls is empty, parse the tool call out of content and lift it into structured tool_calls. Below is the adapter I added to Hermes; the same idea applies to any OpenAI-compatible agent framework.

Hook point (in the response-normalization path of the agent loop), gated to the local/custom provider so it can never reinterpret a genuine text answer from a real OpenAI/Anthropic backend:

# After normalizing the assistant message, before deciding finish_reason:
if agent.provider == "custom" and not getattr(assistant_message, "tool_calls", None):
    from agent.text_tool_calls import recover_tool_calls_from_text
    if recover_tool_calls_from_text(assistant_message):
        finish_reason = assistant_message.finish_reason  # now "tool_calls"
Adapter module (text_tool_calls.py) — parses the text formats local models emit

It recognises the wrapper-tag JSON forms FL emits (<tools>, <tool_call>, <response>, <function_call>, <tool_calls> — single object or list, arguments or parameters), the Qwen XML function syntax (<function=NAME><parameter=KEY>VALUE</parameter></function>), and bare/fenced JSON with name + arguments. On success it populates tool_calls, strips the call text from content, and flips finish_reason to "tool_calls".

import json, re, uuid
from typing import Any

_WRAPPER_TAGS = ("tool_call", "tool_calls", "tools", "response", "function_call")
_TAG_RE = re.compile(r"<(%s)\b[^>]*>(.*?)</\1>" % "|".join(_WRAPPER_TAGS), re.DOTALL | re.IGNORECASE)
_XML_FUNC_RE = re.compile(r"<function=([^>\s]+)\s*>(.*?)</function>", re.DOTALL | re.IGNORECASE)
_XML_PARAM_RE = re.compile(r"<parameter=([^>\s]+)\s*>(.*?)</parameter>", re.DOTALL | re.IGNORECASE)
_FENCE_RE = re.compile(r"^```[a-zA-Z0-9_-]*\s*|\s*```$")

def _coerce_scalar(v: str) -> Any:
    v = v.strip()
    try: return json.loads(v)
    except Exception: return v

def _strip_fences(t: str) -> str:
    return _FENCE_RE.sub("", t.strip()).strip()

def _normalise_json_call(obj: Any):
    if not isinstance(obj, dict): return None
    name = obj.get("name") or obj.get("function")
    if not isinstance(name, str) or not name.strip(): return None
    args = obj.get("arguments")
    if args is None: args = obj.get("parameters", {})
    args_json = args if isinstance(args, str) else json.dumps(args or {})
    return {"name": name.strip(), "arguments": args_json}

def parse_text_tool_calls(content):
    if not content or ("<" not in content and "{" not in content): return [], []
    calls, spans = [], []
    for fm in _XML_FUNC_RE.finditer(content):
        params = {}
        for pm in _XML_PARAM_RE.finditer(fm.group(2)):
            params[pm.group(1).strip()] = _coerce_scalar(pm.group(2))
        calls.append({"name": fm.group(1).strip(), "arguments": json.dumps(params)})
        spans.append((fm.start(), fm.end()))
    if calls: return calls, spans
    for tm in _TAG_RE.finditer(content):
        try: parsed = json.loads(_strip_fences(tm.group(2)))
        except Exception: continue
        items = parsed if isinstance(parsed, list) else [parsed]
        found = [c for c in (_normalise_json_call(o) for o in items) if c]
        if found:
            calls.extend(found); spans.append((tm.start(), tm.end()))
    if calls: return calls, spans
    for candidate in (content.strip(), _strip_fences(content)):
        try: parsed = json.loads(candidate)
        except Exception: continue
        items = parsed if isinstance(parsed, list) else [parsed]
        found = [c for c in (_normalise_json_call(o) for o in items) if c]
        if found: return found, [(0, len(content))]
    return [], []

def recover_tool_calls_from_text(normalized) -> bool:
    if getattr(normalized, "tool_calls", None): return False
    content = getattr(normalized, "content", None)
    if not isinstance(content, str) or not content.strip(): return False
    calls, spans = parse_text_tool_calls(content)
    if not calls: return False
    # build_tool_call is the agent's helper for an OpenAI-shaped tool call
    normalized.tool_calls = [build_tool_call(id=f"call_{uuid.uuid4().hex[:24]}",
                                             name=c["name"], arguments=c["arguments"])
                             for c in calls]
    normalized.content = None
    normalized.finish_reason = "tool_calls"
    return True

Why gate to the local provider: the recovery must never run against a real OpenAI/OpenRouter/Anthropic response, or a genuine text answer containing JSON/angle-brackets could be misread as a tool call. Gating by provider makes it a safe, targeted shim.


Verified result

With FL running the CPU model and the adapter in place, the agent completes a full tool-calling turn:

$ hermes --yolo -t file -z 'Create a file at /tmp/proof.txt containing exactly: FL_TOOLCALL_OK'
The file /tmp/proof.txt has been successfully created with the text "FL_TOOLCALL_OK".
$ cat /tmp/proof.txt
FL_TOOLCALL_OK

Plain chat works too (17 * 351).

Ask

The adapter is only necessary because of Bug 1 (#874). If FL surfaces structured tool_calls under tool_choice="auto" (server-side), any OpenAI-compatible agent — Hermes, LangChain, the OpenAI SDK, custom loops — works against Foundry Local with zero client-side patching. Happy to help test a fix on the same setup.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with #874 and the /v1/chat/completions handling for tool_choice="auto", comparing its response with "required" using the supplied curl reproduction. Trace the server-side tool-call extraction path and verify completion when auto returns populated tool_calls with finish_reason="tool_calls", without requiring the Hermes client-side adapter.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
ai, api, backend-api-design
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.