Self-Improving Bayesian Sentiment Analysis for Twitter
danzambonini.com
danzambonini.com
Semi-supervised learning is a good idea in this type of situation, given that unlabeled samples are far more abundant than labeled samples, but there are gotchas to watch out for. In general, SSL helps when your model of the data is correct and hurts when it is not.
Here's an example of what can go wrong in this particular application: let's say the word 'better' is mildly positive, but when it appears in high-confidence samples, it's usually because it appears together with the words 'business' and 'bureau', as in "I just reported Company X to the Better Business Bureau", i.e., strongly negative. This means that the new self-training samples containing the word better will all be negative, which will bias the corpus until eventually 'better' is treated as a strongly negative feature.
Occasional random human spot-checks of the high-confidence classifications would be useful :-) Also, self-training gives diminishing returns in accuracy, whereas the possibility for craziness remains, so turning it off after a while might be best.
A survey of semi-supervised learning: http://www.cs.wisc.edu/~jerryzhu/pub/ssl_survey.pdf
Basically, the problem is that you can't make a closed system learn from itself, without any outside feedback. The information has to come from somewhere.
It's a bit like someone giving you two Chinese phrases and their translation (without you knowing any Chinese beforehand), and then leaving you to translate a whole book. You will start guessing, based on what you already know, and by the end you'll have arrived to a (totally incorrect) interpretation of what you think each ideogram means.
Context is everything in natural language processing, and by dropping all context the problem becomes harder to solve.
For a classifier that's a less useful approach, but I think single words is too narrow. 3-grams is probably the sweet spot for something like this.
given that unlabeled samples are far more abundant than labeled samples
It is also related to Transfer learning. This paper provides a good background http://www.stanford.edu/~hllee/icml07-selftaughtlearning.pdfIt contains a nice graphic describing difference between Supervised, Semi-Supervised, Transfer and Self-Taught learning.
Since twitter data consists of short texts. Unlabeled data from sources other than twitter could also be used, e.g. Google Buzz.
http://itunes.apple.com/WebObjects/MZStore.woa/wa/viewiTunes...
The explanation of the 'naive' part of Naive Bayes isn't quite right, though. Throwing out the possibility of the animal being human based on the datapoint "four legs" is orthogonal to naivety. A more sophisticated Naive Bayes system could reject classifying an animal as human based on having four legs, and conversely a non-naive system i.e. one that used joint probabilities of the features, might not be more likely to.
I guess the most rational way to do that would be to express and calculate the conditional probabilities with some statistical distance, like the # of standard deviations. So 100k examples of humans without a single one having the "four legs" feature would make that a very strong indicator of being non-human. And that'd work just as well with a naive algorithm as any other.
Is there a name for this?