aws-samples / aws-samples/amazon-textract-serverless-large-scale-document-processing

Some PDFs not generating any outputs

Aperta
#15 1 commento 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
337
Fork
159
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

Whilst on the whole this is working really well, we have found that some PDFs that we upload into the textractpipelinestack-documentsbucketxxx bucket do not get processed completely.

They do seem to still be being processed by textract as the charges are still being applied to our account, but the outputs do not load into the bucket as expected within the code.

I cannot seem to find any error messages suggesting that it isn't working.

I can provide examples of documents if needed. We have a number where they work completely fine, but others simply don't output.

I have raised a Support ticket as well, but they suggested that I raise an issue on here as well.

We want to use textract quite heavily when we move into product build / production, but we are concerned that some documents aren't being picked up.

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia riproducendo il problema con PDF che producono e non producono output nel bucket textractpipelinestack-documentsbucketxxx, quindi traccia il percorso di elaborazione e output di Amazon Textract descritto nel repository. Il lavoro è completato quando la causa degli output mancanti è stata identificata e i documenti interessati producono gli output attesi nel bucket, con un percorso di errore osservabile per gli errori.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
aws, python
Ambito
cloud
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.