Participation At Scale Can Repair The Public Square https://www.noemamag.com/participation-at-scale-can-repair-t...
Polis: Scaling deliberation by mapping high dimensional opinion spaces https://scholar.google.com/scholar?q=Polis:+Scaling+delibera...
Restricting clustering to 2-5 groups impacts group aware/informed consensus and comment routing https://github.com/compdemocracy/polis/issues/1289
That said, something like hdbscan doesn’t suffer from this problem.
(n.b. Celebi's, note usage of Lab / notHSL, respect Cartesian / polar nature of inputs / outputs, and ask for high-K, 128 is what we went with but it's arbitrary. Can get away with as few as 32 if you're ex. Doing brand color from favicon)
For each pixel instead of a single color value it generated k mean color values, using an online algorithm. These were then combined to produce the final pixel color.
The idea was that a pixel might have several distinct contributions (ie from different light sources for example), but due to the random sampling used in path tracing the variance of sample values is usually large.
The value k then was chosen based on scene complexity. There was also a memory trade-off of course, as memory usage was linear in k.
What does online mean here?
In my case the algorithm uswd would use the first k samples as the initial means, and would then find which of the k means were closest to the current sample, and update that mean[2].
Given that in parh tracing one would typically use a fairly large number of samples per pixel relative to k, this approach did a reasonable job of approximating the k means.
If you want accuracy at an order of magnitude more compute, you can use something like DBSCAN.
Built as part of a larger carpet based localisation project [2]
1: https://nbviewer.org/github/tim-fan/carpet_color_classificat...
2: https://github.com/tim-fan/carpet_localisation/wiki/Carpet-L...
So my question is... could you elaborate?
Imagine that you have a dataset, where you think there are likely meaningful clusters, but you don't know how many, especially where it's many-dimensioned.
If you pick a k that is too small, you lump unrelated points together.
If k is too large, your meaningful clusters will be fragmented/overfitted.
There are some algorithms that try to estimate the number of clusters or try to find the k with the best fit to the data to make up for this.
Easier just to deliberately overshoot (with a too high k) and then merge any clusters with too much overlap.
Assume you collect some kindergartners and top NBA players into a room and collect their heights. Now say you pass these to two hapless grad students and ask them to perform K-means clustering.
Suppose one of the grad students knew the composition of the people you measured and can guess these height should clump into 2 nice clusters. The other student who doesn't know the composition of the class - what should they guess K to be?
I understood the GP's comment to refer to the state of the second grad student. How useful is K-means clustering without knowing K in advance?
There are several heuristics for this. Googling I see that the elbow method, average sillhouette method and gap statistic method is the most used.
I think you could play around with your own heuristics as well. Simple KDE plots showing the amount of peaks. Maybe, say the variance between clusters should be greater than the variance inside any cluster could maybe work. (Edit: this seems to be the main point of the average sillhouette method).
k-means assumes these hard parts of the problem are taken care of and offers a trivial solution to the rest. Thanks for the help, I'll cluster it myself.
It’s honestly fine for just finding key differences like a principal component for light storytelling. They don’t need to be distinct clusters