aws-samples / aws-samples/amazon-textract-textractor
access extra functionality like in amazon-textract-response-parser
- Dominant language
- Jupyter Notebook
- Stars
- 493
- Forks
- 163
- PR merge metrics
- No merged PRs in 30d
Description
In https://github.com/aws-samples/amazon-textract-response-parser/blob/master/src-python/README.md I can see several features that I'd like to access from amazon-textract-textractor.
Specifically:
- Order blocks (WORDS, LINES, TABLE, KEY_VALUE_SET) by geometry y-axis
- Page orientation in degrees
- Merge or link tables across pages
- Add OCR confidence score to KEY and VALUE
- Getting the table headers (not mentioned in amazon-textract-response-parser) but available in Textract
Is this possible?
Contributor guide
Research direction
Start with src-python/README.md in the linked amazon-textract-response-parser project and compare its listed features with this repository's current capabilities. Define the scope for geometry ordering, page orientation, cross-page tables, confidence scores, and table headers; done means the requested functionality is available in amazon-textract-textractor.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, python
- Domain
- backend, data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100