cloudflare / cloudflare/cloudflare-os

Multi-turn agent conversations 400 on several Workers AI models (content-parts arrays / null content rejected by per-model schemas)

Open
#54 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
TypeScript
Stars
9.9k
Forks
1.2k
Avg merge
1d 20h
Merged PRs (30d)
107

Description

## Summary

When a Workers AI model is added as a user model (REST transport, "Other Cloudflare Workers AI..."), simple chat works, but agent runs fail with `Error: 400 status code (no body)` as soon as the tool-call loop begins. The cause is that pi sends OpenAI-standard multi-turn payloads (content-parts arrays, and `content: null` on assistant messages carrying `tool_calls`) to `/ai/v1/chat/completions`, and the per-model input schemas of several Workers AI models only accept plain-string, non-null `content`.

This hits free-plan users especially hard: the two suggested models (Kimi K2.7 Code, GLM 5.2) are Paid-only (the 403/5035 from #21), so free-plan users must enter a model ID manually — and 3 of the 4 obvious free choices reject the agent's multi-turn shape.

## Repro

1. `pnpm run-local` (commit aedcda8, Linux), sign up, add model "Other Cloudflare Workers AI..." with `@cf/openai/gpt-oss-120b` + account ID + Workers AI API token (free plan).
2. Chat "Hello" → works.
3. Ask "Make a tic tac toe game." → agent creates the gadget, then the run dies with `Error: 400 status code (no body)`.

The response body (captured with a debug fetch injected into the model transport):

```
AiError: Bad input: Error: oneOf at '/' not met, 0 matches: required properties at '/' are 'prompt',
Type mismatch of '/messages/0/content', 'array' not in 'string',
Type mismatch of '/messages/1/content', 'array' not in 'string',
Type mismatch of '/messages/2/content', 'array' not in 'string',
Type mismatch of '/messages/3/content', 'array' not in 'string',
Type mismatch of '/messages/4/content', 'string' not in 'null',
required properties at '/messages/4' are 'role,content', ...
```

Interestingly, a single user message with a content-parts array passes on gpt-oss-120b — it's the multi-turn mix (multiple array messages + a `content: null` assistant message with `tool_calls`) that fails, so simple chat works and agent runs die.

## Per-model probe

Sending the same OpenAI-standard multi-turn shape (system/user as content-parts arrays, assistant `content: null` + `tool_calls`, `role: "tool"` result, trailing user message) directly to `/ai/v1/chat/completions`:

| model | result |
|---|---|
| @cf/openai/gpt-oss-120b | 400 (schema) |
| @cf/meta/llama-3.3-70b-instruct-fp8-fast | 400 (schema) |
| @cf/zai-org/glm-4.7-flash | 200 OK |
| @cf/qwen/qwen2.5-coder-32b-instruct | 400 (schema) |

So the schemas are per-model. The curated suggested models presumably accept the full shape (glm-4.7-flash, same family as GLM 5.2, does) — which would be why paid users never see this. On the free tier, the only combination that works out of the box appears to be glm-4.7-flash, and nothing surfaces that to the user.

## Suggested fix

Since the per-model schemas are outside this repo's control, normalizing at the model transport seems like the right layer (right where the `NOTE(binding-transport)` comment in `packages/workshop-backend/src/ai-models.ts` already anticipates injecting a fetch shim). This 15-line normalization made the full agent loop (gadget build → runtime-error round-trip → working gadget) complete on gpt-oss-120b:

```ts
fetch: async (url, init) => {
if (String(url).includes("/ai/v1/chat/completions") && typeof init?.body === "string") {
const body = JSON.parse(init.body);
for (const m of body.messages ?? []) {
if (Array.isArray(m.content)) {
m.content = m.content.map(p => typeof p === "string" ? p : p?.text ?? "").join("\n");
} else if (m.content == null) {
m.content = "";
}
}
init = { ...init, body: JSON.stringify(body) };
}
return fetch(url, init);
}
```

(Only safe for text-only models; a per-model compat flag in pi would be the cleaner variant. Related class of issue: #36, where the Ollama slot sends `role: "developer"` that OpenAI-compatible endpoints reject.)

Happy to provide more logs if useful. Write-up of the full investigation (Japanese): https://toming.app/tech/2026/08/cloudflare-os-workers-ai/

Contributor guide

Open the contributing guide

Research direction

Start in packages/workshop-backend/src/ai-models.ts at the NOTE(binding-transport) comment and run pnpm run-local with @cf/openai/gpt-oss-120b to reproduce the multi-turn 400. Compare the transport request with the per-model schema behavior, then verify that the full gadget build and runtime-error agent loop completes for affected models without breaking compatible models.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, backend-api-design
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
74/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.