google / google/langextract

Does it support extracting structured information from the results recognized by OCR

Open
#121 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
38.6k
Forks
2.7k
Avg merge
3d 15h
Merged PRs (30d)
3

Description

Suppose I now use a certain OCR model to extract MarkDown files from product information PDFS or some contract form-based PDFS. Can langextract reasonably parse structured information from md files? If possible, I think this will be a really great project

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.