NIPS 2014 papers
cs.stanford.edu
cs.stanford.edu
Brown seems to have picked up on linear algebra. "Vector", "matrix", "tensor" and "decomposition" all get consistently labeled brown, as do "eigenvalues", "orthogonal" and "sparse".
The rest are not as useful. Black almost always has "number", "set", "tree" and "random", but little else. Purple at times seems to signify topic modeling, but also contains "neural" and "feedforward". Blue seems to be the stats topic, containing "Bayes", "regression", "gaussian", and markov processes. But it also contains random words like "university" and "international".
Overall, very interesting. I wonder if these topics would be even better defined with a higher setting of k.
In addition to adjusting k, another change that might be interesting would be to include also previous years' papers in the model estimation. Changes in component (topic) weights year-over-year could perhaps reveal something about the topics, or the papers.
There are many machine learning libraries that have good implementations of LDA (e.g. Gensim), so it should be "relatively" straightforward to create the topics and clustering based on the abstracts of the papers.
But one of the listed papers is also by Kapathy ("Deep Fragment Embeddings for Bidirectional Image Sentence Mapping"), and I think this might be what nl is complimenting as being done quickly.
The Karpathy paper, too.
I love the cross-modal work that's going on at the moment.