NVIDIA / NVIDIA/Personal-AI-Router

[Bug]: OpenAI-compatible tool calls are returned as plain text when routed through PAIR

Abierto
#94 1 comentario 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

bug
Lenguaje dominante
Go
Estrellas
1.4k
Forks
250
Merge medio
23 h 27 min
PR fusionados (30 d)
1

Descripción

PAIR version or commit

0.1.1-463

Affected component

Ollama proxy

Environment

OS: Ubuntu 26.04.1 LTS on model host; Linux on NVIDIA DGX Spark PAIR client
Architecture: x86_64 model host; arm64 PAIR client
GPU and driver: AMD Radeon 8060S / Ryzen AI Max+ 395 on model host; NVIDIA GB10 on PAIR client
Engine and version: Ollama 0.33.2
Model: qwen3.6:27b
Cluster size: 4 nodes

Steps to reproduce
  1. Configure a PAIR cluster with an Ollama node serving qwen3.6:27b.

  2. Send an OpenAI-compatible POST request to:
    /v1/chat/completions

    The request contains multiple OpenAI function/tool definitions and asks the model to work on a task for which it should call the kanban_show function.

  3. Send the request through the PAIR Ollama proxy.

  4. Send the exact same JSON request body directly to the Ollama server hosting qwen3.6:27b, bypassing PAIR.

  5. Compare finish_reason, message.content, and message.tool_calls in the two responses.

Repeated control test:

  • Direct to Ollama: 3/3 requests returned a native tool_calls object.
  • Through PAIR: 3/3 requests returned finish_reason: stop with no tool_calls object. The intended function call instead appeared as ordinary assistant content.

The behavior is reproducible with curl and does not require an agent framework.

Expected behavior

When an OpenAI-compatible chat completion request containing tool definitions is routed through PAIR, native tool-calling behavior should be preserved.

For this request, the response should contain message.tool_calls with a call to kanban_show, and finish_reason should be tool_calls, matching the response obtained when the identical request is sent directly to the Ollama server.

Actual behavior

When the identical request is sent through the PAIR Ollama proxy, the native OpenAI-compatible tool call is not returned.

In repeated testing, PAIR returned:

  • finish_reason: stop
  • message.tool_calls: null / absent
  • The intended kanban_show call embedded in message.content as ordinary text.

Three consecutive PAIR requests produced textual representations such as:

  1. Python-like syntax:
    kanban_show(task_id="t_65ba4229")

  2. Bracket syntax:
    [kanban_show][0]

  3. XML-like syntax:
    <invoke>kanban_show(task_id="t_65ba4229")</invoke>

By contrast, three consecutive requests using the identical JSON request body sent directly to Ollama returned native OpenAI-compatible tool calls:

finish_reason: tool_calls

with message.tool_calls containing a function call to kanban_show.

The same JSON payload succeeds when sent directly to the Ollama server on the model host, so the failure is introduced only when the request is routed through PAIR.

This also prevents OpenAI-compatible agent/tool frameworks from recognizing and executing the requested function when the request is routed through PAIR.

Sanitized logs or screenshots

Confirmations
  • I searched existing issues for duplicates.
  • This is not a security vulnerability.
  • I agree to follow the Code of Conduct.

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Comienza en el punto de entrada /v1/chat/completions del proxy de Ollama y compara cómo se gestionan las solicitudes y las respuestas cuando hay definiciones de herramientas. Reprodúcelo con la solicitud curl documentada contra PAIR y directamente contra Ollama, y verifica después que PAIR conserve message.tool_calls y finish_reason: tool_calls en lugar de devolver la llamada como texto.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
go, ollama
Área
api, backend
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Activo
Claridad
Bastante claro
Aptitud para principiantes
55/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.