aboutcode-org / aboutcode-org/scancode-analyzer

An Analysis of the Existing Rules

Aperta
#12 0 commenti 0 reazioni 1 assegnatario Rivendicata da @AyanSinhaMahapatra Vedi su GitHub
Lingua principale
Python
Stelle
4
Fork
4
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

1. Load in all the Rules/License-Texts in a Pandas Dataframe (and one without the texts for faster loading).

2. Analysis on the following contexts

a. Minimium Coverage
b. Relevance Scores
c. License Type [“unknown”, “other-copyleft", "permissive" etc]
d. Rule Length
e. Most Used Words across License Types (and their variance)
f. Rules vs License Texts

3. Attempt Clustering/Visualizations of Rule-Texts and Various Inaccuracies in License Scan Results. (T-SNE)

This would be used later, as all the Scan Results will also be analyzed in these contexts, and this data will be required for creating new rules automatically, automate License Type Detection using NLP, etc tasks.

Format - Jupyter Notebook with related Functions loaded from Modules.

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.