docling-project / docling-project/docling

Setup and processing steps for VLM

Open
#2,713 3 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
2d 21h
Merged PRs (30d)
84

Description

### Question

...

What are the steps for using a VLM model for information extraction from a pdf. What should be done first to download and make available required models and set the path to the models. Is there any tool within the docling repository to download the required models and set the path variable ? How is a extraction pipelien defined when we have pre-existing fields as per a json structure to be extracted. The pdf is of mixed format where there are scanned areas of text and also plain textual areas.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.