docling-project / docling-project/docling

OCR is consistently missing spaces

Open
#2,887 6 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
3d 4h
Merged PRs (30d)
98

Description

### Bug
When ingesting a document without embedded text the OCR is concatenating a significant amount of the text together

### Steps to reproduce
run document converter on the following public pdf
https://tinyurl.com/48krwcjw

### Docling version
2.68

### Python version
3.9.5

Contributor guide

Open the contributing guide

Research direction

Start at the document converter entry point and run it against the linked public PDF using Docling 2.68 and Python 3.9.5, focusing on the OCR path for documents without embedded text. Compare the extracted text with the PDF to identify where spaces are lost; done means the OCR output preserves the document's word spacing.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.