aws-samples / aws-samples/amazon-textract-serverless-large-scale-document-processing
Files with more than 200 pages are not completely extracted
- 主要言語
- Python
- スター
- 337
- フォーク
- 159
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
I noticed a consistent problem with larger files(200+ pages)
PDF files with more than 200 pages are never completed extracted. The files either remain unprocessed(no analysis folder created) or just partially extracted(40-50 pages) even after 1-2 days. Files within 100 pages are extracted within 1-5 mins.
The process is not bombarded with many large files, the problem is same even if I upload one large file(200 pages) in day.
Please let me know if I am missing something for larger files.
Thank you
コントリビューションガイド
調査の方向性
まず、200ページを超える単一のPDFで問題を再現し、その後、正常に完了する100ページ未満のファイルと比較します。ドキュメント処理のワークフローを追跡して、抽出が停止する理由、または分析フォルダーが作成されない理由を特定します。大きなドキュメント全体が抽出され、その分析フォルダーが作成されれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- aws, python
- 領域
- backend, cloud
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 35/100