Topic mining with LDA and Kmeans and interactive clustering in Python
ahmedbesbes.com
ahmedbesbes.com
One thought I had was that I would be careful using KMeans on tSNE transformed data. From my understanding, while tSNE preserves local relationships between data point, it (heavily) distorts long-distance relationships. What this means is that (if I get this right), you have almost arbitrary clustering results if your k is chosen such that it is not equal to the number of local clusters (a thing you don't know in advance). So clustering on distances between data points in tSNE coordinates seems to be a risky way to draw conclusions. I am not saying this is wrong, but if what I am saying is correct (please point out if I got something wrong), then your analysis may perhaps work better if you did KMeans on say the final 5-dimensional PCA space or something like that, and leave tsne to visualize the data (and perhaps label the data points by the clustering on PCAs?). Do you agree / does that make sense?
kmeans = kmeans_model.fit(vz)
kmeans_clusters = kmeans.predict(vz)
this is easy enough to do, I'm wondering what exactly is your definition of "interesting" ?
here is an example of taking tweets that have 'diabetes' and correlating topics with counties. https://www.linkedin.com/in/karl-dailey-02557b65/treasury/po...
The original topics were actually better, but I was asked to adjust the language to wipe out common language across all tweets. You can still see interesting things going on though. North East: Charity, Hospitals, Research. South: Koolaid, sweat tea. etc.
I had also done topic modeling on customer surveys for Comcast (I cant show them), but the topics identified 3 key features that lead to low customer satisfaction.
I have also used LDA on grocery purchases (for mere fun), and it worked out really great as well.
A more interesting topic would be one that is understandable but was not obvious, even to the people intimately involved in data's subject matter. These are harder to discover, but they do exist - and I have never seen LDA surface them.
If it cant give you what you are looking for, maybe you are asking a fish to climb a tree
negative reviews: http://pastebin.com/vkGn7ZLH
Amazon Echo, roughly ~34k reviews. https://en.wikipedia.org/wiki/Non-negative_matrix_factorizat...
It's surprisingly useful for finding interesting trends in reviews.
The thing is, reviews, messages, comments, etc, in general, do tend to revolve around some central topic(s). For example, this comment right now is about LDA. It's not totally random.
Then when presented with a new article on any subject create the vocabulary and frequencies for that article, strip out the words stripped previously and then do a match against all subjects.
Very easy to parallelize, very good results (surprisingly good, for such a simple algorithm in fact).
I have a dumb question about Bokeh ... if I use RStudio I can create some markdown, export to PDF, send around to team/clients. Typically I can't do anything interactive without hosting on a shiny server or referencing some server-side magic.
If I use Bokeh, is it as simple as creating a Jupyter notebook, saving to HTML, and then anyone can see all the D3 magic? Or are there any drawbacks like file size, compatibility, browsers not letting you open due to security or whatnot? If any pointers, tutorials, examples anyone has would love to hear them, look me up if it's not deemed on-topic.
http://rstudio.github.io/crosstalk/
Edit: Note that you can knit rMarkdown to HTML as well, which is what enables most of this stuff.
for next project maybe I'll try both ways, with ggplot and rmarkdown to HTML, and with Bokeh in python.
The ggplot will just knit to PNGs, would be interesting if ggplot could output D3 with intelligent rollovers, maybe eventually some scriptable animations.
One could use ggplotly but I don't want to post anything to, or host anything on a server, has to work 100% offline.