microsoft / microsoft/markitdown

Extraction is not in markdown

Open
#206 5 comments 8 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
186k
Forks
13.7k
Avg merge
1d 4h
Merged PRs (30d)
49

Description

I tried to extract the contents of pdf. But it is extracting as plain text, not as markdown. Am I missing any parameter?

from markitdown import MarkItDown
md = MarkItDown()

result = md.convert("microsoft_report.pdf")
print(result.text_content)

output_file = "output.md"
with open(output_file, "w", encoding="utf-8") as file:
file.write(result.text_content)
print(f"Markdown content has been written to {output_file}")
Image

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the report through MarkItDown().convert("microsoft_report.pdf") and inspect result.text_content before it is written to output.md. Determine whether the returned content contains the expected Markdown structure; done means the PDF conversion behavior is corrected or clearly documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
tooling
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.