aboutcode-org / aboutcode-org/scancode-toolkit

ScanCode fails to scan large unrar test file

Abierto
#712 2 comentarios 0 reacciones 0 asignados Ver en GitHub
complex fixed pending review timeout
Lenguaje dominante
Python
Estrellas
2.6k
Forks
791
Merge medio
1 d 12 h
PR fusionados (30 d)
5

Descripción

Reported by @msrb in https://github.com/fabric8-analytics/fabric8-analytics-worker/issues/172

The culprit is in the `nulls.txt` file in https://pypi.python.org/pypi/py-unrar2/0.99.6 : this is a ~100MB test text file with zeroes that is the thing that ScanCode chokes on. Actually this file contains 100M times the '0' character.
```
$ python -c "print len(set(open('nulls.txt', 'rb').read()))"
1
```

There are a couple ways to deal with this that come to mind:
- only scan up to a certain size for large files (configurable with an option) and return a warning
- compute the number of unique characters in a file (or start of a file) and use this with a threshold to classify the file as data. Do not scan "data" files. This could be refined by computing the entropy.
- if the file is "text" and the size above a certain threshold, classify the file as data. Do not scan "data" files.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.