aws-samples / aws-samples/amazon-textract-serverless-large-scale-document-processing

How to get particular page Form-data.

オープン
#41 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

主要言語
Python
スター
337
フォーク
159
PR マージ指標
30日以内にマージされた PR はありません

説明

I Have added the entire PDF file to S3Bucket via amazons3Client.PutObjectAsync method and started the analysis by textractClient.StartDocumentTextDetectionAsync method.
while getting the response by textractClient.GetDocumentTextDetectionAsync method I can see all pages data and could able to segregate the Raw/Line data.
My problem is, how can I get the FormData and TableData for a particular page(say I need FormData only for page no 3). Kindly advise on this.

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

まず、issue に記載されている S3 PutObjectAsync、Textract StartDocumentTextDetectionAsync、GetDocumentTextDetectionAsync の呼び出しと、リポジトリのドキュメント処理のエントリーポイントを確認します。ページ固有のフォームデータとテーブルデータがどのように表現されるかを確認し、その後、サポートされているアプローチと、3 ページ目の明確で完全な完了例を文書化します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
aws, python
領域
documentation
issue の種類
ドキュメント
難易度
3/5
見積もり時間
1〜2日
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。