aws-samples / aws-samples/amazon-comprehend-semi-structured-documents-annotation-tools
issues and slowness of tool
- Dominant language
- Python
- Stars
- 24
- Forks
- 15
- PR merge metrics
- No merged PRs in 30d
Description
We are scaling out a data labelling project on this tool, it is going very well and we are excited to see the NER comprehend results once it has the rich positional data to make predictions. There are some roadblocks we have in though:
can only tag one twice
Can only label 1 page at a time, would like to remove "Read only",
order of tasks/pages,
converting ui - template to have input fields,
zoom in on pdf,
performance is slow in general
We feel like all of these are addressable, however because some of this code is minified and essentially not open source, this adds some risk into trying to fork this and start working on it.
Would love to talk to someone on a call about all of this. I opened a support ticket to this effect in our aws account, but they directed me to put an issue in here. Please advise, I can provide contact info if we want to set up a call.
Contributor guide
Research direction
The issue lists several separate concerns: repeated tagging, page access and ordering, editable UI templates, PDF zoom, and general slowness. Start by reproducing each concern in the annotation tool, then split confirmed problems into focused issues with specific completion criteria and relevant code or tests identified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- frontend, machine-learning, performance, tooling
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100