Latent Dirichlet Allocation on Tweets
wellecks.wordpress.com
wellecks.wordpress.com
[1] edit: Because I'm not sure if I just made that phrase up or if it came from one of your papers, the idea that ML libraries that take declarative model descriptions are great, but what's even better is if we also have an imperative API that can dynamically generate those declarative specs for us, even based on train-time inputs, so we can essentially "program" the structure of a model but still benefit from keeping everything generalized and declarative at the base.
It would be interesting to extend the LDA model to include a temporal variable. Never got around to doing it, but it seems like it would work well for social media data.
[0] http://blog.dc.esri.com/2013/04/18/the-evolution-of-discussi...
[1] http://blog.dc.esri.com/files/2013/04/topic-distribution2.pn...
http://josephmisiti.github.io/using-latent-dirichlet-allocat...
I've read that LDA doesn't work well on short documents. Your approach of concatenating all tweets for a user appears to work quite well. One other technique I've seen is to concatenate multiple tweets together that contain the same hashtag.
One of our intern students at 99designs did some work on applying LDA to classify graphic design tasks:
http://99designs.com.au/tech-blog/blog/2014/01/22/Swiftly-Ma...
.. you might find it interesting. :)