docling-project / docling-project/docling

docling is not providing better results for arabic language

Open
#1,421 8 comments 0 reactions 2 assignees Claimed by @cau-git View on GitHub
question
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
3d 4h
Merged PRs (30d)
98

Description

I have been working on arabic pdf.what if arabic pdf consist of scanned images ?.At this stage , we have to classify wheather the page is machine readable text or scanned image.Do we have a to way figure out wheather page has machine readable text or scanned image .
If we are able to figure that out, then we can apply ocr pipeline to that particular page.
I would be grateful, if anyone provide me with required configurations to get better results for arabic language.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.