aws-samples / aws-samples/amazon-textract-textractor

Modification to the Document entity from response is not captured when using .to_trp2()

Open
#120 1 comment 1 reaction 0 assignees View on GitHub
enhancement
Dominant language
Jupyter Notebook
Stars
493
Forks
163
PR merge metrics
No merged PRs in 30d

Description

The conversion to trp2 is based on using the initial response. This does not capture the any modifications made to the entities like OCR post-processing or correction or deletion of entities. A proper converter needs to be implemented to make the library usable for post-processing in-place modifications.

This would allow workflows along the lines of:
```
document.pages[1].key_values = {key: value + '_edited' for key, value in document.pages[1].key_values}
document.export("document.json")
```

We also need to add utilities function such as
- [ ] Merging tables
- [ ] Adding new keys
- [ ] Adding new queries output

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.