aws-samples / aws-samples/amazon-textract-serverless-large-scale-document-processing
Files with more than 200 pages are not completely extracted
Open
- Dominant language
- Python
- Stars
- 337
- Forks
- 159
- PR merge metrics
- No merged PRs in 30d
Description
I noticed a consistent problem with larger files(200+ pages)
PDF files with more than 200 pages are never completed extracted. The files either remain unprocessed(no analysis folder created) or just partially extracted(40-50 pages) even after 1-2 days. Files within 100 pages are extracted within 1-5 mins.
The process is not bombarded with many large files, the problem is same even if I upload one large file(200 pages) in day.
Please let me know if I am missing something for larger files.
Thank you
Contributor guide
Assessment
This issue has not been assessed yet.