docling-project / docling-project/docling
OCR is consistently missing spaces
Open
bug
- Dominant language
- Python
- Stars
- 66.4k
- Forks
- 4.8k
- Avg merge
- 3d 4h
- Merged PRs (30d)
- 98
Description
### Bug
When ingesting a document without embedded text the OCR is concatenating a significant amount of the text together
### Steps to reproduce
run document converter on the following public pdf
https://tinyurl.com/48krwcjw
### Docling version
2.68
### Python version
3.9.5
Contributor guide
Research direction
Start at the document converter entry point and run it against the linked public PDF using Docling 2.68 and Python 3.9.5, focusing on the OCR path for documents without embedded text. Compare the extracted text with the PDF to identify where spaces are lost; done means the OCR output preserves the document's word spacing.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100