Feature clarification: AI provider enhancements — thinking mode, vision image input, and API host normalization
- Dominant language
- Go
- Stars
- 15.7k
- Forks
- 1.4k
- Avg merge
- 3d 8h
- Merged PRs (30d)
- 7
Description
### Background
We run a self-hosted Apache Answer as a community FAQ site, with an
OpenAI-compatible LLM gateway (SenseNova). While integrating the AI feature we
hit three concrete gaps. Following @LinkinStars's suggestion on #1594, this
issue clarifies the proposed feature set before the PR is resubmitted.
### Proposed enhancements
**1. API host normalization (bugfix)**
The admin AI settings accept an "API host", but the model-listing endpoint
concatenates `api_host + "/v1/models"` while the chat client auto-appends
`/v1`. A base URL that already contains `/v1` — the normal convention for
OpenAI-compatible gateways (SenseNova, SiliconFlow, OneAPI relays, vLLM,
SGLang...) — therefore produces `/v1/v1/models` and fails with the gateway's
`NOT_FOUND` error, which is passed through to the UI verbatim.
Proposal: a shared `NormalizeAPIHost()` helper (trim whitespace/trailing
slashes, append `/v1` when missing, keep `/v1beta/*` endpoints intact) used by
both code paths, plus concise upstream error summaries instead of echoing raw
response bodies.
**2. Per-provider thinking mode switch**
Reasoning-capable models are increasingly common (DeepSeek V4, Qwen3,
SenseNova, ...). Proposal: an admin switch per AI provider that injects the
OpenAI-compatible `enable_thinking: true` flag into chat completion requests.
The existing `reasoning_content` streaming/rendering path already displays the
thought process, so no front-end change is required for display.
**3. Per-provider vision (image input) switch**
Proposal: an admin switch enabling image attachments in AI conversations —
up to 4 PNG/JPEG/WebP images (≤4 MB decoded each) sent as base64 data URLs or
HTTPS links, converted to MultiContent parts. The switch is exposed via site
info (`ai_vision_enabled`) so the chat input only shows the attach button for
vision-capable providers. History records keep a textual placeholder instead
of persisting image data, so no DB schema change is needed.
### Scope / non-goals
- No new provider protocols (OpenAI-compatible only); Anthropic removed from
the default provider list since the backend speaks the OpenAI protocol.
- No multi-provider routing or per-user model selection.
- No persistence of image attachments in conversation history.
### Related
- #1594 (will be closed and resubmitted referencing this issue).
- Regarding the security reminder in #1594: the flagged constant was the base64
of a standard 1×1-pixel PNG placeholder used in a unit test, not a credential.
To avoid any misreading, the test now constructs the image at runtime from its
signature bytes and the source contains no base64 blob.
Happy to hear feedback on the scope before resubmitting. Thanks!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by tracing the model-listing endpoint and chat client, then follow provider settings into chat completion requests, site info, the chat input, and conversation history. Confirm the three proposed provider enhancements and their stated limits are covered, with concise upstream errors and no persisted image data; the existing reasoning-content rendering path should remain usable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- go
- Domain
- ai, api, backend, frontend
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100