microsoft / microsoft/markitdown

markitdown-ocr: PPTX converter emits literal \n sequences in Markdown output

Open
#2,010 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
186k
Forks
13.7k
Avg merge
1d 4h
Merged PRs (30d)
49

Description

PptxConverterWithOCR uses "\\n" instead of "\n" in multiple output paths.
As a result, converted PPTX files contain literal backslash-n sequences rather
than Markdown line breaks.

The OCR PPTX tests currently encode this malformed output as expected behavior.

Additionally, the optional LLM caption path imports ._llm_caption, but the
module exists in markitdown.converters, not in markitdown_ocr.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the PptxConverterWithOCR implementation and the OCR PPTX tests mentioned in the issue. Check each output path for newline handling, then inspect the optional LLM caption import and the markitdown.converters module location. Update the affected expectations and run the OCR PPTX tests to confirm Markdown line breaks and caption imports work.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, content
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.