Development of Data Extractor from Document(PDF, image files for now) , with OCR and PDF reader
Open
Task
- Dominant language
- Python
- Stars
- 8
- Forks
- 2
- Avg merge
- 14h 34m
- Merged PRs (30d)
- 46
Description
Title of ticket:
#### Description
Document on S3 need to be pulled based on the 'disconnected` mechanism using REDIS Stream ~ similar to Dedupe , Stiching etc.
#### Dependencies
Are there any dependencies?
#### DOD
- [ ] List the items that need to be complete for this ticket to be considered done
- [ ]
- [ ]
- [ ]
- [ ]
Contributor guide
Research direction
No files, tests, or entry points are named. Start by tracing the existing disconnected mechanism and Redis Stream flows used by Dedupe and Stitching, then determine how S3 documents and OCR/PDF reading should fit. Done requires explicit acceptance criteria for pulling documents and extracting data from PDF and image files.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, python, redis
- Domain
- backend, data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100