microsoft / microsoft/markitdown
markitdown-ocr: PPTX converter emits literal \n sequences in Markdown output
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 186k
- Forks
- 13.7k
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 49
Description
PptxConverterWithOCR uses "\\n" instead of "\n" in multiple output paths.
As a result, converted PPTX files contain literal backslash-n sequences rather
than Markdown line breaks.
The OCR PPTX tests currently encode this malformed output as expected behavior.
Additionally, the optional LLM caption path imports ._llm_caption, but the
module exists in markitdown.converters, not in markitdown_ocr.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the PptxConverterWithOCR implementation and the OCR PPTX tests mentioned in the issue. Check each output path for newline handling, then inspect the optional LLM caption import and the markitdown.converters module location. Update the affected expectations and run the OCR PPTX tests to confirm Markdown line breaks and caption imports work.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend, content
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 68/100