How we analysed 70M comments on the Guardian website
theguardian.com
theguardian.com
I'd be concerned for the analysis itself that using the blocking of a comment as a proxy for the comment being abusive contains significant risk of confirmation bias - moderators may be more inclined to moderate articles by authors with traditionally female names, or those authors may be more inclined to report comments, or indeed they may be more inclined to write articles on more controversial topics. I can't think of a better automated approach to the analysis, and I find the outputs pretty believable, but I'd be wary of regarding the study as evidence.
At the same time it doesn't moderate comments that are simple insults if they agree with its angle on other topics, such as its aggressively institutionalised dislike of Labour leader Jeremy Corbyn.
This seems to be part of an attempt to link commenters who don't agree with the official line with the worst excesses of online trolling.
It's a rather nasty move, and trying to pretend it's a scientific analysis doesn't make the G look good.
>It is also important to say that it is true that the research is only valid if you accept a) that the Guardian's community standards are fit for purpose, and b) that the moderators are reasonably skilled at applying them. Some readers may not accept either of these things, and our research is not designed to justify these assumptions.
https://www.theguardian.com/technology/2016/apr/12/how-we-an...
Given 70 million comments I might hope for something more, like perhaps a proposition as to how to better automatically catch the bad comments without having to have moderators at all, via some sort of lexical analysis of the good ones versus the bad ones from this manually sorted set.
Would the Guardian be willing to release the data set for others to analyse?
It's useful to me to see some attempt to quantify this and to get some numbers that show what the balance is.
The obvious example that springs to mind would be Germaine Greer's article in the sports section which was nominally about Caster Semenya but mostly just attacked trans women, who she's got a particular dislike for. Little chance a male Guardian journalist would get that through editorial, no chance it'd be allowed in the comments, and since she's not even a sports writer normally there's no non-inflammatory articles from her to counterbalance it.
"Our list of authors contains the approximately 12,000 individuals"
so guardian is more like a blog aggregator ?
>In the future, we would like to explore the words used in the comments, using standard and bespoke natural language processing algorithms.
There's some interesting things from a tech perspective - use of Postgres, Perl scripts etc. That was fun to find out.
Unfortunately the Guardian is in regular battle over their comments, particularly in respect to how they handle comment regarding Israel - they are frequently accused of suppressing anything 'pro' (by deleting entire commenter accounts as breaches of their rules) and retaining some quite alarming comments (on Israel and other subjects) by classifying as free speech and retaining despite complaints - see [1] for 68 'recommends' for killing Tony Blair being kept.
This looks like an attempt to bypass some serious issues by intentionally addressing something completely tangential. A classic diversionary tactic.
That's why I think you (or anyone else) is unlikely to get access to the data.
Which is a pity, because it could help clear up exactly these issues of claimed bias they are accused of, and help determine if it exists or not. - something useful.
[1] https://ukmediawatch.org/2011/12/31/cif-commenter-suggests-t...
a) a comment was abusive
b) AND it was posted under a given article
c) THEN it was abusive towards the author of the article
Based on this assumption they find a correlation with articles being written by women and minorities, and high rates of abuse (a.k.a. moderated comments). The assumption is a very big one to make. People in comments may just be hurling abuse at each other. They might be trolling each other rather than the author of the article, so their abuse may have nothing to do with the sex or minority status of the author.
Did they check that this is not the case? They report that they did check that the correlation was, to their opinion, statistically significant [1] but they don't give any details of their test for this.
In general, it's true that The Guardian has a big problem with comments. Eyballing it, most of their comments are dross that doesn't contribute anything to any conversation, even if they're not abusive. They need to attract a better class of commenter. They'd do well to take a page out of HN's book and implement a self-moderation system.
Also, allowing editing of comments for a brief period might help improve their quality. Sometimes people might regret pressing "post" but have no way to retract their own comment.
[1] https://discussion.theguardian.com/comment-permalink/7221868...
"“The Guardian, once a standard bearer of quality journalism now contains football journalists so in love with Mourinho it makes me sad. This is just the latest in an incredible long campaign for the despicable one to join the club of Matt Busby and Jimmy Murphy. I am astonished that the editor of the paper allows this dross to be published. You are a disgrace to the profession.”"
I know nothing of this soccer caper but aside from the last sentence, it seems like a fairly typical rant. But why do they publish Mourinho-loving dross? :)
If they genuinely want to clean up their comments section, eliminate the professional trolls. e.g. On any politics story the same faces appear, seemingly plucked from university debating teams. It doesn't matter what ridiculous position is being argued as long as they cheered for their side and smeared the other mob. Most voters have leanings but don't have a "team" per se and simply want good policy.