aws-samples / aws-samples/amazon-textract-serverless-large-scale-document-processing

Files with more than 200 pages are not completely extracted

Aperta
#22 0 commenti 1 reazione 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
337
Fork
159
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

I noticed a consistent problem with larger files(200+ pages)
PDF files with more than 200 pages are never completed extracted. The files either remain unprocessed(no analysis folder created) or just partially extracted(40-50 pages) even after 1-2 days. Files within 100 pages are extracted within 1-5 mins.
The process is not bombarded with many large files, the problem is same even if I upload one large file(200 pages) in day.
Please let me know if I am missing something for larger files.
Thank you

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.