docling-project / docling-project/docling
Question: Is it possible to configure images that are extracted from other images?
Open
question
- Dominant language
- Python
- Stars
- 66.4k
- Forks
- 4.8k
- Avg merge
- 2d 21h
- Merged PRs (30d)
- 84
Description
### Question
Hi all! Thanks for the great work on docling. Quick question: Is it possible to configure a minimum size (or similar) for when images are extracted out of other images in the docling pipeline? When images are extracted from a PDF I pass to docling, I end up with parts of the images in different places in the document. I'd like to avoid this if possible, and keep only the "entire" image (or at least be able to configure this). Note: I'm not running OCR in the pipeline - only text and image extraction.
Contributor guide
Assessment
This issue has not been assessed yet.