aboutcode-org / aboutcode-org/scancode-toolkit

ScanCode fails to scan large unrar test file

Aberta
#712 2 comentários 0 reações 0 responsáveis Ver no GitHub
complex fixed pending review timeout
Linguagem predominante
Python
Estrelas
2.6k
Forks
791
Merge médio
1d 12h
PRs com merge (30d)
5

Descrição

Reported by @msrb in https://github.com/fabric8-analytics/fabric8-analytics-worker/issues/172

The culprit is in the `nulls.txt` file in https://pypi.python.org/pypi/py-unrar2/0.99.6 : this is a ~100MB test text file with zeroes that is the thing that ScanCode chokes on. Actually this file contains 100M times the '0' character.
```
$ python -c "print len(set(open('nulls.txt', 'rb').read()))"
1
```

There are a couple ways to deal with this that come to mind:
- only scan up to a certain size for large files (configurable with an option) and return a warning
- compute the number of unique characters in a file (or start of a file) and use this with a threshold to classify the file as data. Do not scan "data" files. This could be refined by computing the entropy.
- if the file is "text" and the size above a certain threshold, classify the file as data. Do not scan "data" files.

Guia de contribuição

Abrir o guia de contribuição

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.