aboutcode-org / aboutcode-org/scancode-toolkit

Consider dropping pdfminer and use XPDF to extract text from PDF

オープン
#1,865 コメント 5 件 リアクション 0 件 担当者 0 名 GitHub で見る
new feature
主要言語
Python
スター
2.6k
フォーク
791
平均マージ
1日 12時間
マージ済み PR(30日)
5

説明

# Short Description
pdfminer is both slow and has been the source of more than a few issues in the past. Xpdf is C code and os-specific but the pdftotext command may be just enough of what we need:
http://www.xpdfreader.com/download.html

It comes with pre-built command line tools for Linux, Windows and Mac

## Possible Labels
- new feature

## Select Category
- Enhancement [x]
- Add License/Copyright []
- Scan Feature []
- Packaging []
- Documentation []
- Expand Support []
- Other []

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。