aws-samples / aws-samples/amazon-textract-serverless-large-scale-document-processing

Files with more than 200 pages are not completely extracted

オープン
#22 コメント 0 件 リアクション 1 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
337
フォーク
159
PR マージ指標
30日以内にマージされた PR はありません

説明

I noticed a consistent problem with larger files(200+ pages)
PDF files with more than 200 pages are never completed extracted. The files either remain unprocessed(no analysis folder created) or just partially extracted(40-50 pages) even after 1-2 days. Files within 100 pages are extracted within 1-5 mins.
The process is not bombarded with many large files, the problem is same even if I upload one large file(200 pages) in day.
Please let me know if I am missing something for larger files.
Thank you

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

まず、200ページを超える単一のPDFで問題を再現し、その後、正常に完了する100ページ未満のファイルと比較します。ドキュメント処理のワークフローを追跡して、抽出が停止する理由、または分析フォルダーが作成されない理由を特定します。大きなドキュメント全体が抽出され、その分析フォルダーが作成されれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
aws, python
領域
backend, cloud
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。