aboutcode-org / aboutcode-org/scancode-toolkit

Consider dropping pdfminer and use XPDF to extract text from PDF

Aperta
#1,865 5 commenti 0 reazioni 0 assegnatari Vedi su GitHub
new feature
Lingua principale
Python
Stelle
2.6k
Fork
791
Merge medio
1g 12h
PR unite (30g)
5

Descrizione

# Short Description
pdfminer is both slow and has been the source of more than a few issues in the past. Xpdf is C code and os-specific but the pdftotext command may be just enough of what we need:
http://www.xpdfreader.com/download.html

It comes with pre-built command line tools for Linux, Windows and Mac

## Possible Labels
- new feature

## Select Category
- Enhancement [x]
- Add License/Copyright []
- Scan Feature []
- Packaging []
- Documentation []
- Expand Support []
- Other []

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.