aws-samples / aws-samples/amazon-textract-serverless-large-scale-document-processing
Some PDFs not generating any outputs
- Lingua principale
- Python
- Stelle
- 337
- Fork
- 159
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Whilst on the whole this is working really well, we have found that some PDFs that we upload into the textractpipelinestack-documentsbucketxxx bucket do not get processed completely.
They do seem to still be being processed by textract as the charges are still being applied to our account, but the outputs do not load into the bucket as expected within the code.
I cannot seem to find any error messages suggesting that it isn't working.
I can provide examples of documents if needed. We have a number where they work completely fine, but others simply don't output.
I have raised a Support ticket as well, but they suggested that I raise an issue on here as well.
We want to use textract quite heavily when we move into product build / production, but we are concerned that some documents aren't being picked up.
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia riproducendo il problema con PDF che producono e non producono output nel bucket textractpipelinestack-documentsbucketxxx, quindi traccia il percorso di elaborazione e output di Amazon Textract descritto nel repository. Il lavoro è completato quando la causa degli output mancanti è stata identificata e i documenti interessati producono gli output attesi nel bucket, con un percorso di errore osservabile per gli errori.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- aws, python
- Ambito
- cloud
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Ferma
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100