Clustering related stories
blog.getprismatic.com
blog.getprismatic.com
We focus solely on sports and our classifiers(supervised) reduce the scope first to the sport and then to the specific team before we apply any sort of clustering (k-means, LDA, etc). That allows us to reduce the vocabulary to what is mostly a list of named entities for the sport/team and key words such as 'injury', 'quarterback', etc. With a significantly reduced vocabulary, even algorithms such as Hierarchical LDA work surprisingly well.
years ago i got my hands wet with some rudimentary text classification and grouping techniques. it was when google news came out, and it was fluffier than i like my news to be (it has since been tuned a lot better). so i wrote my own.
i wrote up my (admittedly naive and simplistic - but effective) methods here:
http://www.informit.com/articles/printerfriendly.aspx?p=3988...
In retrospect it's pretty ugly, but it worked pretty well. I really wanted to implement named entities and n-grams but never got around to it. I'm glad you guys did :)
I didn't account for names entities or n-grams in the feature vector though. That's a very interesting idea.
@mattdeboard - what algorithm did you use to count the occurrence and size of clusters?
1. scalability: does your system ingest multiple documents in parallel? if so, how often do you observe over-segmentation, if any?
2. thresholds: how did you set the thresholds at various parts of the systems?