docling-project / docling-project/docling-parse

docling is not providing better results for arabic language

Open
#118 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
333
Forks
80
Avg merge
2d 4h
Merged PRs (30d)
12

Description

I have been working on arabic pdf.what if arabic pdf consist of scanned images ?.At this stage , we have to classify wheather the page is machine readable text or scanned image.Do we have a to way figure out wheather page has machine readable text or scanned image .
If we are able to figure that out, then we can apply ocr pipeline to that particular page.
I would be grateful, if anyone provide me with required configurations to get better results for arabic language.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.