Introducing Streaming K-Means in Spark MLlib 1.2
databricks.com
databricks.com
http://en.wikipedia.org/wiki/Jenks_natural_breaks_optimizati...
Do you have any advice or ideas for automatically picking count of clusters in an unknown data set?
You could do that here too, just have a range of k. If only there were a streaming leave-one-out cross validation for k-means to complement this approach ...
(it is possible to do this in a streaming style, see LWPR)
K-Means without preset K.