Heat map of homophobic and racist tweets
users.humboldt.edu
users.humboldt.edu
"Pet Peeve #208: Geographic profile maps which are basically just population maps."
When zoomed out the 100th meridian population divide is as obvious as it could be. So At the global level normalization is pretty much not in effect. This might likely be caused by the not uniform distribution and size of counties in the US.
A better approach might be binning in equal area regions, with a threshold for low sample sizes.
ie: maybe there's just way more tweeting in Austin.
[edit: I'm apparently wrong. @ronaldx points out below a statement that suggests this is already happening and is why Orange County doesn't shine brighter]
The three hottest counties in New Mexico and Idaho have populations under 34000 combined. One of them is just 2900 people.
The number of users tweeting hate and how connected they are to others is probably a more interesting map.
I would disgree in that we have a geographic claim that, for example, Nebraska and Iowa supposedly have spectacular difference in hate levels, like order of magnitude, yet simultaneously claims that Salt Lake City and San Fran are equally gay-friendly. And on a micro level, the hyper-segregated highly urban area to my east has no racist tweets at all, yet farmville east of the mississippi is supposedly a hotbed of racism but magically it all disappears west of the mississippi. I've lived in both Wisconsin and Alabama and anyone is seriously claiming the racial climate is basically identical? LOL.
Nope, we're being trolled here. I do agree with you that more data, or more analysis, would make it more interesting.
It could be titled something like "The Apparent Relative Visibility of Hate in Tweets at the County Level". "The Geography of Hate" isn't such a far jump from there, especially if you are just trying to show that the tweets are coming from all over.
So, given that this is a rate-based map, the XKCD doesn't apply
"Orange County, California has the highest absolute number of tweets mentioning many of the slurs, but because of its significant overall Twitter activity, such hateful tweets are less prominent and therefore do not appear as prominently on our map."
=========================================
The data behind this map is based on every geocoded tweet in the United States from June 2012 - April 2013 containing one of the 'hate words'. This equated to over 150,000 tweets and was drawn from the DOLLY project based at the University of Kentucky. Because algorithmic sentiment analysis would automatically classify any tweet containing 'hate words' as "negative," this project relied upon the HSU students to read the entirety of tweet and classify it as positive, neutral or negative based on a predefined rubric. Only those tweets that were identified by human readers as negative were used in this analysis.
To produce the map all tweets containing each 'hate word' were aggregated to the county level and normalized by the total twitter traffic in each county. Counties were reduced to their centroids and assigned a weight derived from this normalization process. This was used to generate a heat map that demonstrates the variability in the frequency of hateful tweets relative to all tweets over space. Where there is a larger proportion of negative tweets referencing a particular 'hate word' the region appears red on the map, where the proportion is moderate, the word was used less (although still more than the national average) and appears a pale blue on the map. Areas without shading indicate places that have a lower proportion of negative tweets relative to the national average.
The numbers that appear in the map during a mouse hover indicate the total number of hateful tweets and number of unique users sending them in each county.
==========================================
EDIT: The mouse overs don't appear to work very well in Chrome or Firefox, but from the one or two times I was able to see some numbers it appears that each red circle may be a dozen or less tweets. Also, the hot zones dissipate significantly the further you zoom in, so without any statistics or numbers it's difficult to draw conclusions.
A very interesting experiment, but given that the data is only normalized by Twitter traffic (non-response bias) this is in no way indicative of the actual distribution of racism.
I wonder how well a Bayesian classifier would work if the this was used as a training set. If it worked relatively well, there's no reason why you couldn't create a live version of the map.
Something like http://aworldoftweets.frogdesign.com/ maybe?
Consider using millions of training examples (vs. thousands). This was done as part of the "distant supervision" Twitter sentiment technique. What this means is that tweets with positive emoticons were labeled as positive sentiment, and negative emoticons were labeled as having negative sentiment. Emoticons were stripped before training. This system got 80% accuracy.
http://cs.wmich.edu/~tllake/fileshare/TwitterDistantSupervis...
It's also unclear what criteria was used to define negative. If a reference to Lonely Island's "no homo" song gets flagged as negative for example, that may not fit in most people's definition of hate speech.
[Edit: I should place extra emphasis on the word "roughly," in my OP as well.]
An interesting distinction you've made there.
It would be interesting to identify people who have posted just a few hateful messages—perhaps few enough that they can get away with it in their local social context. This may more sensitive to occult haters.
Something like this: 1. Identify individual twitter users as being hateful in a particular category. For instance, user A uses the word "chink" 3x and "fag" 5x in 100 tweets, so he gets added to the "chink" and "fag" categories. Play with these threshold to see what makes sense. 2. Divide the # of hateful users in each category by the total # of users in that location. Allocation of users to location can be done proportional to the # of tweets they make from each location.
Cool project :).
It makes a big difference if red is 4-5 tweets (per what time period), or 4-5 thousand.
Put differently: If county A has 100 tweets over the time period and 1 of them are racist it will appear identical to county B which has 1000 tweets of which 10 are racist over the same time period.
I think we're being trolled. The one guy in the north woods of wisconsin who uses twitter used the "n word" once but no one in the entire 4 million person metro Milwaukee area has. Hmm. The Milwaukee MSA has 0.5% of the population of the entire USA and about 90% of the black population of Wisconsin, so it should have included 750 tweets based on supposed sample size yet none of them in troll happy twitter were *-ist in any form. Unlikely. I think we're being trolled, perhaps as a sociology test of gullibility, and being fed data on "decade of founding of city" or something equally irrelevant. Or perhaps just random. Or data intentionally designed to make the east look worse than the west, although it seems weird to merge Wyoming in culturally with San Franciso or Berkley in with Colorado.
A better complaint about the word homophobic is that its meaning is dependant on one understanding "homo" to mean homosexual. Otherwise it means one who fears same-ness. Most homophobes actually, obviously, fear difference.
That said, homophobia, in cultures like North America, is very much based on fear. It's not a simple sort of fear, but it is wrapped up in how gay people challenge gender roles, and the fear that men have of being less than dominant. People who aspire to socially dominant roles, who secretly have same-sex feelings, are often the worst oppressors. Before you condemn the term, I suggest you read up on it.
Because algorithmic sentiment analysis would automatically classify any tweet containing 'hate words' as "negative," this project relied upon the HSU students to read the entirety of tweet and classify it as positive, neutral or negative based on a predefined rubric. Only those tweets that were identified by human readers as negative were used in this analysis.
If you have 50% people of the racial group you are sampling racist tweets from living in a particular state, you should see less racist tweets by %.
Thanks for an interesting chart.
What did you do to cut down on lying human classifiers? Do you give them an incentive? Did you have them vote as a group on the classification of a sentiment?
Whites aren't allowed to say the N word, and Anita can say the S word http://www.youtube.com/watch?v=Qy6wo2wpT2k&t=0m45s. But maybe your human classifiers can handle this problem too? I hope they got a good grade.
Emphasis mine.
I don't understand how normalization is supposed to protect privacy though.