docling-project / docling-project/docling-serve

How to improve OCR recognition accuracy

Open
#552 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.8k
Forks
340
Avg merge
6d 11h
Merged PRs (30d)
8

Description

**Question**
I want to use EasyOCR, but the results are not very satisfactory, and I'm looking for a solution.

First of all, my test images are mainly in Japanese and English. I added DOCLING_SERVE_ARTIFACTS_PATH="", but the recognition quality is still not very good. However, the results are very good in the demo at https://www.jaided.ai/easyocr/. I'd like to know how to solve this issue.

Of course, if I can recognize the test images better, I don't necessarily have to use EasyOCR.

I would be extremely honored to receive your help.

test image:
![Image](https://github.com/user-attachments/assets/867be82f-b37c-4192-bbc3-bcacb1415507)

parsed image:
Image

easy ocr demo:
Image

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the provided test image, parsed image, and EasyOCR demo comparison, along with the DOCLING_SERVE_ARTIFACTS_PATH setting. Reproduce the reported Japanese and English recognition results and determine whether the issue is configuration, model selection, or an OCR integration problem; done means a confirmed cause and a specific, reproducible improvement.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.