Monitoring mood in the UK using Twitter
rawkes.com
rawkes.com
I realize that this isn't the key focus of your paper, but we've found that sampling and human analysis/tagging is far more accurate at judging the sentiment around a brand, company or topic.
In the context of this study, I found that it was impossible to accurately infer 'sentiment' of a single tweet or person (not just because of sarcasm and other nuances). However, when you take the average of a group (wisdom of the crowd) then the results are much more promising. A trends noticed across thousands of users is also more interesting than the potentially unreliable sentiment of a single person.
In this case, I suppose it is definitely just the words that are being analysed – not true 'sentiment.' I wouldn't rule it out as inaccurate though, it just depends what you're looking for and how you use the results. Compared to other sentiment data-sets, the ANEW approach seems more more detailed (the original scorings are created from human tagging).
I do agree though, that automated approaches can be inaccurate if you're looking for fine-level analysis.
It's also worth pointing out that the data-set used in this study to gauge 'sentiment' is based on firm psychology and actually infers much more than just perceived happiness.
Effectively, yes. It's called the Wisdom of the Crowd: http://en.wikipedia.org/wiki/Wisdom_of_the_crowd
To rule out artefacts you need to work out a) what you're looking for, and b) whether it is backed up by anything else. For example, in this study my findings are backed up by other, different studies. The findings also correlate with key public events.
Being able to infer true sentiment is not what is claimed here. Instead, you're able to infer 'sentiment' trends that are backed up in some way by other studies and research.
I'm 100% sure that these approaches aren't perfect, however they are proving useful. For example, one group of people are using a very similar approach to take average 'sentiment' on Twitter and use it to predict stock market fluctuations 3 days in advance. It works and it's proven not to be fluke. Something is in the results, however inaccurate a single tweet is.
I'm well aware of the concept of the wisdom of the crowd but if your incoming data is noise then your ability to build any kind of aggregate analysis on top of it is going to be nil & Post-hoc analysis means you can find only things you already know are there. ANEW/AFINN etcetc are basically the white flag to any kind of meaningful automated analysis of tweets and resorting to simply dumb word counts instead. Yes, you can capture a broad pattern but the shape and strength of that? The contours are an artifact of the list you use, it's utterly arbitrary. Throw a few more random phrases on the list, assign them some arbitrary values and presto, new results! If an tweet goes round twitter in a minute with hundreds of thousands of RTs this kind of analysis will miss it completely unless it's lucky enough to be preloaded with the right dictionary. To work in this kind of context an algorithm has to be able to trim its own sails.
There's nothing particularly wrong with your work, I'm just sick of the cycles wasted attempting to do an extreme version of the problems that NLP already struggles with.
I'm surprised the corpus of words is so small and that there is no attempt to provide context. It just wasn't what I was expecting. I was expecting more of a markov chain kind of setup.
In Swedish, for example, there are lots of negative words that are trendy to use in a way that is extremely positive, rather like "wicked!" in English slang.
Context is sacrificed with the ANEW approach, though that's not to say that accuracy and results are compromised. It just means that you need to look at the output in a different way, after understanding the limitations of the input.
In ANEW (the data-set used), words that imply multiple meanings are likely to receive more neutral values (it's all coded by humans).
The SentiWordNet data-set is probably a little more like what you're expecting: http://sentiwordnet.isti.cnr.it/
Good news for the country perhaps, but it ruined my visualisation plans...
https://code.google.com/p/sasa-tool/
and as you state, even the most negative tended to come out as neutral at best. However, even a basis analysis of the hashtags used seemed to show that most messages were positive. It seems that Twitter as a whole leans somewhat to the left.
My original plan had been to create a heatmap (similar to one I made here: http://heatmapdemo.alastair.is/), with red and blue 'clouds' to indicate which parts of the country were happy and which were angry as the inauguration went on, but the data just doesn't seem to be out there.
As for the sentiment analysis… I'm unable to release the data-set as it's owned by a university in Florida and only available to students. A quick Google on sentiment analysis, or 'Affective Norms for English Words' will come up with useful things. :)