annoviko / annoviko/pyclustering

Documentation for using CURE with large datasets

Ouverte
#633 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub
Proposal
Langage dominant
Python
Étoiles
1.2k
Forks
262
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

I did not find any documentation for how to use CURE with large (2M+) datasets. Simply using the cure algorithm as is defined in cure.py is not feasible since the building of the queue and kd-tree itself will take significant time.

I noticed that there is a functionality for random sampling, added in response to a feature to include random sampling for CURE for this very reason. However I am not clear on how to use it.

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.