MiniMax-AI / MiniMax-AI/cli

text chat: expose image input (enable multimodal M3, incl. multi-image)

Open
#224 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
TypeScript
Stars
2.2k
Forks
181
Avg merge
9h 4m
Merged PRs (30d)
13

Description

Summary

mmx text chat uses the multimodal MiniMax-M3 model by default, but the CLI exposes no image input flag. Users cannot send an image (let alone multiple images) to M3 through text chat without hand-writing a base64 messages JSON file. Since M3 is multimodal, this is a significant hidden capability.

Current behavior

  • mmx text chat only accepts text: --message, --messages-file, --system. There is no --image flag.
  • parseMessages (src/commands/text/chat.ts) passes content through as string | ContentBlock[], so image blocks can reach the API via --messages-file — but the CLI does no image handling (no path→base64 conversion, unlike vision describe).
  • For single-image description there is mmx vision describe (which hits /v1/coding_plan/vlm, single-image only).
  • No CLI path exists for multi-image input (compare/diff/joint analysis of 2+ images in one call), even though M3 supports it.

Expected behavior

A first-class image input on text chat, e.g.:

# single image
mmx text chat --model MiniMax-M3 --image ./photo.jpg --message "What breed is this dog?"

# multiple images (repeatable)
mmx text chat --model MiniMax-M3 \
  --image ./before.png --image ./after.png \
  --message "List every visual difference between these two."

The flag should accept local paths / http(s) URLs and auto base64-encode them (reusing toDataUri from src/utils/image.ts), then inject them as image content blocks alongside the text message.

Evidence — multi-image already works via M3

I verified that M3 accepts multiple images in one call through mmx text chat --messages-file. Example (CN region, API key auth):

node -e '
  const fs = require("fs");
  const img = (p) => ({ type: "image", source: { type: "base64", media_type: "image/png", data: fs.readFileSync(p).toString("base64") } });
  fs.writeFileSync("/tmp/m.json", JSON.stringify([{
    role: "user",
    content: [
      { type: "text", text: "I am giving you TWO images. Describe one detail unique to each." },
      img("/tmp/a.png"), img("/tmp/b.png")
    ]
  }]));
'
mmx text chat --model MiniMax-M3 --messages-file /tmp/m.json --non-interactive --quiet
# → M3 correctly describes both images and distinguishes them
Format gotcha worth surfacing

mmx text chat posts to the Anthropic /messages endpoint (chatEndpoint returns ${baseUrl}/anthropic/v1/messages), not the OpenAI /chat/completions endpoint. So image content blocks must use the Anthropic shape:

// ❌ OpenAI shape — rejected: "unsupported content type 'image_url'"
{ "type": "image_url", "image_url": { "url": "data:..." } }

// ✅ Anthropic shape — works
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "<base64>" } }

The OpenAI-compatible image format documented at https://platform.minimaxi.com/docs/api-reference/text-openai-api (type: "image_url") does not work through this CLI's text chat, because the CLI routes to the Anthropic endpoint. This mismatch took a while to debug — a --image flag that auto-formats correctly (or at least a doc note) would help a lot.

Suggested implementation

  1. Add a repeatable --image <path-or-url> flag to text chat.
  2. In parseMessages / the chat body builder, convert each --image via toDataUri, then append { type: "image", source: { type: "base64", media_type, data } } blocks to the user message's content (converting content from string to array when images are present).
  3. When --image is present, default --model to MiniMax-M3 if not set.
  4. Optionally reuse the same --image flag on a future vision subcommand for multi-image, since M3's chat path strictly supersedes the single-image /vlm endpoint for multi-image use cases.

Happy to open a PR if this design sounds reasonable.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in src/commands/text/chat.ts, especially parseMessages and the chat body builder, then inspect src/utils/image.ts and the existing vision describe image handling. Add repeatable local-path or URL image input that produces Anthropic image blocks, supports multiple images, and defaults to MiniMax-M3 when needed; done means text chat can send one or more images without a messages JSON file.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, cli
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
74/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.