bcgov / bcgov/foi-flow

Programming Tool / Library selection for Data extraction including OCR

Open
#3,695 0 comments 0 reactions 0 assignees View on GitHub
Task
Dominant language
Python
Stars
8
Forks
2
Avg merge
14h 34m
Merged PRs (30d)
46

Description

Title of ticket:

Based on the current performance during document processing for stitching , dedupe using Python, its better to check more light weight , threaded, LOW level language( C++/ C / RUST/) for Data extraction, if nothing works then python! . Need to see proper OCR packages are available for the selected lanuague. Procurement OCR component?

#### Dependencies
Are there any dependencies?

#### DOD
- [ ] List the items that need to be complete for this ticket to be considered done
- [ ]
- [ ]
- [ ]
- [ ]

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are named. Begin by defining the comparison criteria for Python, C++, C, and Rust, then identify available OCR packages and procurement requirements. Done should be a documented language and OCR component choice with dependencies and completed acceptance criteria.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, cpp, python, rust
Domain
computer-vision, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.