docling-project / docling-project/docling

Are docling results supposed to be consistent across runs?

Open
#1,159 1 comment 0 reactions 0 assignees View on GitHub
question triage/close-stale
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
3d 4h
Merged PRs (30d)
95

Description

### Question
For a same [pdf file](https://github.com/user-attachments/files/19229658/paper.pdf), the result converted earlier this morning ([the good one](https://github.com/user-attachments/files/19229659/good.md)) and the one I got this evening ([the bad one](https://github.com/user-attachments/files/19229657/bad.md)) are of completely different quality. Worse still, it seems that I can only get bad results now. The bad one has wrong headers, weird line breaks, and random characters.

Does Docling work in a way that it's supposed to get consistent results? Are there potential parameters to tune when it does not seem to understand the structure of the pdf?

Contributor guide

Open the contributing guide

Research direction

Start by comparing the attached paper.pdf with good.md and bad.md, using the reported differences in headers, line breaks, and random characters as evidence. Done means determining whether the same input should produce consistent results and documenting any supported parameters relevant to structure recognition.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.