docling-project / docling-project/docling

VLM as OCR engine

Open
#2,711 6 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
2d 21h
Merged PRs (30d)
84

Description

I would like to use a VLM as an OCR engine. What I envision that normal text in pdfs is converted using the deterministic backend, while images are described using a VLM. Is this currently possible?

If I understand the documentation correctly, if the VLM pipeline is configured the entire PDF is always sent to the VLM, even if the pdf consists of structured text. Is that correct?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.