microsoft / microsoft/markitdown
ChatGPT OCR results are generated in different languages
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 186k
- Forks
- 13.7k
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 49
Description
The text written in Japanese on the image is translated into English and output.
from markitdown import MarkItDown
from openai import OpenAI
client = OpenAI()
md = MarkItDown(llm_client=client, llm_model="gpt-4o")
result = md.convert("example.jpg") ### Japanese Language Image
print(result.text_content) ### English output
In some cases, the entire document will be in English, while in other cases only part of the document (only the title) will be in English.
Depending on the requirements of your RAG, this may not be desirable, so it is better to be able to specify the output language or to fix it to the original language found in the image.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start from the MarkItDown.convert call using llm_client and llm_model="gpt-4o", then trace how image OCR output language is requested and returned. Define whether completion means preserving the detected source language or accepting an explicit output-language option, and verify the Japanese-image example no longer produces unintended English text.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100