docling-project / docling-project/docling

Question: Is it possible to configure images that are extracted from other images?

Open
#2,289 5 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
2d 21h
Merged PRs (30d)
84

Description

### Question
Hi all! Thanks for the great work on docling. Quick question: Is it possible to configure a minimum size (or similar) for when images are extracted out of other images in the docling pipeline? When images are extracted from a PDF I pass to docling, I end up with parts of the images in different places in the document. I'd like to avoid this if possible, and keep only the "entire" image (or at least be able to configure this). Note: I'm not running OCR in the pipeline - only text and image extraction.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.