docling-project / docling-project/docling

Mixed Arabic-English PDFs produce incorrect OCR/output text

Open
#3,462 2 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
2d 21h
Merged PRs (30d)
84

Description

### Bug
Converting PDFs containing both Arabic and English text, the generated .md output is inconsistent
(Output file on screenshots)

### Steps to reproduce
docling "https://tug.ctan.org/macros/latex/exptl/mem/arabic.pdf"

(arabic.pdf used as example)

### Docling version
Docling version: 2.90.0
Docling Core version: 2.74.0
Docling IBM Models version: 3.13.0
Docling Parse version: 5.9.0
Python: cpython-313 (3.13.13)
Platform: Windows-11-10.0.26200-SP0

### Python version
Python 3.13.13

Image
Image

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.