microsoft / microsoft/markitdown
Can not create markdownfile from bengali pdf
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 186k
- Forks
- 13.7k
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 49
Description
I have downloaded a book from archive in .pdf format. Write the below code:
#md = MarkItDown(use_azure=False) # Important: set this flag
md = MarkItDown(use_azure=False, ocr_mode=True)
<!-- Failed to upload "Bharater-Shilpa-sanskritir.pdf" -->
<!-- Failed to upload "Bharater-Shilpa-sanskritir.pdf" -->
file_path = "E:/NLP/Bengali LLM all/DATA/Bharater-Shilpa-sanskritir.pdf"
result = md.convert(file_path)
# Save to .md file
with open("output.md", "w", encoding="utf-8") as f:
f.write(result.text_content)
print("Markdown saved as output.md")
The output file size is 0 kb.
The input file link is : [(https://archive.org/details/in.ernet.dli.2015.266550/page/n135/mode/2up)]
I have downloaded the pdf version.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the issue with the linked PDF and the shown MarkItDown(...).convert call, then inspect the PDF conversion and OCR path reached by convert. Determine why result.text_content is empty for this Bengali PDF; done means the same input produces non-empty Markdown, with a regression check if the relevant test location is found.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- tooling
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 32/100