dask / dask/dask-ml

More scalable alternative to k-means++ based on sampling

Open
#121 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
951
Forks
262
PR merge metrics
No merged PRs in 30d

Description

I have no direct experience with it but this NIPS 2016 paper looks very interesting and has some theoretical guarantees on the approximation.

Fast and Provably Good Seedings for k-Means
http://papers.nips.cc/paper/6478-fast-and-provably-good-seedings-for-k-means.pdf

Here are some benchmarks:

![image](https://user-images.githubusercontent.com/89061/35327169-3690f30c-00f9-11e8-900e-ee95162f27be.png)

![image](https://user-images.githubusercontent.com/89061/35327200-4ddacf88-00f9-11e8-95af-5c665239c6ae.png)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.