docling-project / docling-project/docling

Backend for PDF OCR

Open
#1,330 2 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
66.4k
Forks
4.8k
Avg merge
2d 21h
Merged PRs (30d)
84

Description

### Question
I am currently implementing a pipeline with Docling Conversion of various formats.

However, I am confused about what backend to use in PDFs and image format documents. What I need is :

1. Better OCR
2. Accurately parse the tables.
3. Able to handle handwritten notes and cursive writing

Currently, I am using Rapid OCR with PyPdfiumDocumentBackend. I am not getting better results with handwritten notes.
...

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.