docling-project / docling-project/docling
Are docling results supposed to be consistent across runs?
- Dominant language
- Python
- Stars
- 66.4k
- Forks
- 4.8k
- Avg merge
- 3d 4h
- Merged PRs (30d)
- 95
Description
### Question
For a same [pdf file](https://github.com/user-attachments/files/19229658/paper.pdf), the result converted earlier this morning ([the good one](https://github.com/user-attachments/files/19229659/good.md)) and the one I got this evening ([the bad one](https://github.com/user-attachments/files/19229657/bad.md)) are of completely different quality. Worse still, it seems that I can only get bad results now. The bad one has wrong headers, weird line breaks, and random characters.
Does Docling work in a way that it's supposed to get consistent results? Are there potential parameters to tune when it does not seem to understand the structure of the pdf?
Contributor guide
Research direction
Start by comparing the attached paper.pdf with good.md and bad.md, using the reported differences in headers, line breaks, and random characters as evidence. Done means determining whether the same input should produce consistent results and documenting any supported parameters relevant to structure recognition.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100