Ask HN: Sentiment Analysis – how to handle biased word list lengths?
I'm implementing a simple sentiment analysis algorithm where the authors of the paper have a word list for positive and negative words and simply count the number of occurrences of each in the analysed document and give it a sentiment score the document with:
sentiment = (#positive_matches - #negative_matches) / (document_word_count)
This is normalising the sentiment score by document length BUT the corpus of negative words is 6 times larger than the positive word corpus (around 300 positive words and 1800 negative words) so by the measure above, the sentiment score will likely be negatively biased since there are more negative words to match than positive words.
How can I correct for the imbalance in the length of the positive vs. negative corpuses?
When I run calculate the above sentiment score, I get around 70% of my 2000 document set with negative sentiment scores BUT there is no a priori reason that my document set should be biased towards the negative and I would expect the true 'unobserved' sentiment of the documents to be approximately symmetrical with around half the documents positive and half negative.
I need to somehow come up with a methodology that results in representative sentiment scores to remove the bias introduced by asymmetrical word lists.
Any thoughts / ideas much appreciated :)