aboutcode-org / aboutcode-org/scancode-toolkit
ScanCode fails to scan large unrar test file
- Lingua principale
- Python
- Stelle
- 2.6k
- Fork
- 791
- Merge medio
- 1g 12h
- PR unite (30g)
- 5
Descrizione
Reported by @msrb in https://github.com/fabric8-analytics/fabric8-analytics-worker/issues/172
The culprit is in the `nulls.txt` file in https://pypi.python.org/pypi/py-unrar2/0.99.6 : this is a ~100MB test text file with zeroes that is the thing that ScanCode chokes on. Actually this file contains 100M times the '0' character.
```
$ python -c "print len(set(open('nulls.txt', 'rb').read()))"
1
```
There are a couple ways to deal with this that come to mind:
- only scan up to a certain size for large files (configurable with an option) and return a warning
- compute the number of unique characters in a file (or start of a file) and use this with a threshold to classify the file as data. Do not scan "data" files. This could be refined by computing the entropy.
- if the file is "text" and the size above a certain threshold, classify the file as data. Do not scan "data" files.
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.