aboutcode-org / aboutcode-org/scancode-toolkit

Consider dropping pdfminer and use XPDF to extract text from PDF

Aberta
#1,865 5 comentários 0 reações 0 responsáveis Ver no GitHub
new feature
Linguagem predominante
Python
Estrelas
2.6k
Forks
791
Merge médio
1d 12h
PRs com merge (30d)
5

Descrição

# Short Description
pdfminer is both slow and has been the source of more than a few issues in the past. Xpdf is C code and os-specific but the pdftotext command may be just enough of what we need:
http://www.xpdfreader.com/download.html

It comes with pre-built command line tools for Linux, Windows and Mac

## Possible Labels
- new feature

## Select Category
- Enhancement [x]
- Add License/Copyright []
- Scan Feature []
- Packaging []
- Documentation []
- Expand Support []
- Other []

Guia de contribuição

Abrir o guia de contribuição

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.