allenai / allenai/mmda

Symbol Scraper cannot correctly parse the tokens in some documents

オープン
#20 コメント 1 件 リアクション 0 件 担当者 1 名 @kyleclo が担当を希望しています GitHub で見る
主要言語
Jupyter Notebook
スター
166
フォーク
19
PR マージ指標
30日以内にマージされた PR はありません

説明

For example, for this one https://www.clearinghouse.net/chDocs/public/JC-DC-0001-0001.pdf,
it will merge two words `McGruder` & `JC-DC-001-001` from line 1 and 2 into a single one `McGruderJC-DC-001-001`.

As such, the extracted tokens boxes are incorrect (grey boxes):

image

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。