aws-samples / aws-samples/amazon-textract-textractor

access extra functionality like in amazon-textract-response-parser

Open
#179 11 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
493
Forks
163
PR merge metrics
No merged PRs in 30d

Description

In https://github.com/aws-samples/amazon-textract-response-parser/blob/master/src-python/README.md I can see several features that I'd like to access from amazon-textract-textractor.

Specifically:
- Order blocks (WORDS, LINES, TABLE, KEY_VALUE_SET) by geometry y-axis
- Page orientation in degrees
- Merge or link tables across pages
- Add OCR confidence score to KEY and VALUE
- Getting the table headers (not mentioned in amazon-textract-response-parser) but available in Textract

Is this possible?

Contributor guide

Open the contributing guide

Research direction

Start with src-python/README.md in the linked amazon-textract-response-parser project and compare its listed features with this repository's current capabilities. Define the scope for geometry ordering, page orientation, cross-page tables, confidence scores, and table headers; done means the requested functionality is available in amazon-textract-textractor.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, python
Domain
backend, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.