annoviko / annoviko/pyclustering
Documentation for using CURE with large datasets
Ouverte
Proposal
- Langage dominant
- Python
- Étoiles
- 1.2k
- Forks
- 262
- Métriques de merge des PR
- Aucune PR mergée en 30 j
Description
I did not find any documentation for how to use CURE with large (2M+) datasets. Simply using the cure algorithm as is defined in cure.py is not feasible since the building of the queue and kd-tree itself will take significant time.
I noticed that there is a functionality for random sampling, added in response to a feature to include random sampling for CURE for this very reason. However I am not clear on how to use it.
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Évaluation
Cette issue n'a pas encore été évaluée.