docling-project / docling-project/docling
Picture description with text context
Open
enhancement
triage/close-stale
- Dominant language
- Python
- Stars
- 66.4k
- Forks
- 4.8k
- Avg merge
- 3d 4h
- Merged PRs (30d)
- 98
Description
### Requested feature
I would like to see an option to provide surrounding text context to image or picture enhancement. I think annotation will be better if VLM has some context in which image is provided. So the model is not getting only the image but for example the section which image is embedded.
Contributor guide
Assessment
This issue has not been assessed yet.