Where is the happiest city in the USA?
onehappybird.com
onehappybird.com
I grew up in Napa so I know something was off here. While adults love it, most kids find it boring or even hate growing up there.
Tweets about wine != happiness. Lack of profanity != happiness. Examples - "The wine sucked" or "Vegas is f*cking awesome!" This is an interesting study of vocabulary, but hardly a measure of happiness.
And that's where there might be a problem, the results may (perhaps partially) reflect a structural difference of 'use of language' between areas.
I didn't see it on the list here, I'm assuming because the population was too small.
http://travel.usatoday.com/destinations/story/2011/04/San-Lu...
How the heck did you get that domain?
In other words, start with my results and then find some data to support them.
> we avoid stemming words, i.e., conflating inflected words with their root form, such as all conjugations of a specic verb. For verbs in particular, by focusing on the most fre- quent words, we obtained scores for those conjuga- tions likely to appear in texts, obviating any need for stemming. Moreover, while we observe stemming works well in some cases for happiness measures, e.g., havg(advance)=6.58, havg(advanced)=6.58, and havg(advances)=6.24, it fails badly in others, e.g., havg(have)=5.82 and havg(had)=4.74; havg(arm)=5.50 and havg(armed)=3.84; and havg(capture)=4.18 and havg(captured)=3.22.
That is something I came across myself when I was working on a similar data mining project. But this was a surprise:
> In terms of methodology, our hedonometer could be improved by incorporating happiness estimates for common n-grams, e.g., 2-grams such as `child abuse' and `sex scandal' as well as negated sentiments such as `not happy'
I can understand ignoring n-grams but a simple no-X or not-X is a necessity. In common parlance, 'Peace' > 'No war' > 'No peace' > 'War'. If you assume the simple happiness values of 'Peace' = 1, 'War' = -1, 'No' = -0.5, then the scores end up: Peace=1, No War=-1.5, No peace=0.5, War=-1. Clearly 'War' is worse than 'No War' yet the scores don't correspond. The correct order should be something like: Peace=1, No War=0.5, No peace=-0.5, War=-1. To make this work, 'No' shouldn't be a fixed negative value like -0.5. It should be -(next-word/2). Then 'No Peace' = -0.5 + 1 = 0.5 and 'No War' = 0.5 + -1 = -0.5.
PS: I typed 'War' so many times in the above sentence that I had to check its spelling to confirm it was a real word. http://en.wikipedia.org/wiki/Semantic_satiation strikes again.
I think it can certainly be used as a "qualitative" analysis, but making it a "measurement of happiness" is overdoing it.
Attempting to use tweeted words as indicators of happiness ignores context.
Analysis based on idividual words might miss a lot of subtle context, especially since English is such an idiomatic language. For instance, focusing on individual words (or even small groups of adjacent words) may lead to mistaking a morbid joke for sadness, while a sarcastic insult might be mistaken for happiness.
Still, this is an ingenious way to measure happiness, and I'm looking forward to reading future research that builds on this work.
I can't imagine “restaurant”, “wine”, and even “cheers” going with people who can't afford any.
Regional happiness may or may not correlate to usage of these "happy words" or "unhappy words", especially across local cultures, circumstances (e.g. before or after Obama's election, different regions will respond differently), technology/twitter adoption, etc.