microsoft / microsoft/markitdown

How get image captioning in docx files?

Open
#1,102 6 comments 3 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
186k
Forks
13.7k
Avg merge
1d 4h
Merged PRs (30d)
49

Description

Hey,
I tried to convert docx with images file to md, but It does not do captioning:

from markitdown import MarkItDown
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")  # my local VLLM host
md = MarkItDown(llm_client=client, llm_model="microsoft/Phi-3.5-vision-instruct")

result = md.convert("file.docx")
print(result.text_content)
# .... ![](data:image/png;base64...) ....

What did I do wrong?

Thank you in advance for your reply!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No repository file or test is named. Start by reproducing the example through MarkItDown.convert("file.docx") with the shown local OpenAI-compatible client, then trace the DOCX conversion behavior. Done would mean that images produce caption text rather than only a data URI, with the expected behavior covered by a test.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
tooling
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.