The Erdős paradox: When a mathematical number and Wikipedia collide
blog.wikimedia.org
blog.wikimedia.org
[1]: http://www.wolframalpha.com/input/?i=1+-+(1+-+12%2F262)%5E12...
Only 17.47% (as of 26 February 2018) of the English Wikipedia's biographies are about women: https://en.wikipedia.org/wiki/Wikipedia:WikiProject_Women_in...
EDIT: I rechecked the quote from the article about the likelihood of deletion and found it to be "in this cohort of Erdős Ones, the pages for women were over four times more likely than men’s to be nominated for speedy deletion." (my italics). FYI.
Wikipedia being a tertiary source may, through unbiased processes, simply be reflecting bias in the secondary sources. Biographies of pre-modern women are far less common than of their male contemporaries, IME; that's less true of my modern and especially extremely current figures, but Wikipedia isn't exclusively focused on current figures.
However if it were 17.47% for current times, that would be a significant bias that needs to be addressed. It still might not hit 50/50 given reduced opportunities for women still today, but it should be moving ever closer.
Talking about p-values in contexts where one has all the facts and the additional fact that one has all the facts suggests a gross misunderstanding of statistical concepts, their utilities, and the circumstances in which they apply.
First of all I at least don't have all the information. If I had it would indeed be meaningless to infer the reason behind the deletion attempts, as I'd know already, but the fact remains that I don't. By this logic trying to see how likely this was to happen assuming articles were just randomly marked for deletion is a perfectly reasonable test.
Second of all, how exactly is this situation different from tossing a coin several times? If I toss a coin 50 times, record the results, and then melt down the coin, surely it'd still be valid to conclude that it was biased if it came up heads 49 times out of 50? Or conversely that it wasn't biased if it came up tails 26 times out of 50?
Arguing that statistics only applies to experiments that can be repeated suggests a gross misunderstanding of statistical concepts, their utility and the circumstances in which they apply.
Maybe I'm misunderstanding Ben's point, but I don't think he's objecting to this as a frequentist. You have correctly calculated that given the small sample size of female candidates, there is about a 10% chance that random deletion of entries would produce a disparity this large.
I think Ben's issue is that since we already know that the deletion process is nonrandom, this tells us nothing we do not already know. The question we're actually interested in is what the factors are that cause deletion requests, not just the ones that correlate with them, and the numerical value you've calculated doesn't give us much insight.
So maybe you could talk more about what you think the calculated p-value actually means, and what we can do with it?
Viewed from that perspective what this p-value means is that merely knowing the proportion of articles that gets marked for deletion is already enough to conclude that it is likely (or at least not impossible) for a group of 21 articles to have at least 2 articles marked for deletion. In that sense the knowledge that of the 21 articles written on women, 2 were marked for deletion doesn't really tell us something that we didn't already know.
Of course that doesn't mean we've exhausted all available knowledge, but you're going to have to look further than just how many articles get deleted.
Various statistical tests exist for taking this into account. With Fisher's exact test [0], we find a $p$-value close to yours:
In [6]: from scipy.stats import fisher_exact, binom
In [7]: binom(12, 12/262).sf(1)
Out[7]: 0.10211478280975775
In [8]: fisher_exact([[2, 10], [10, 240]])[1]
Out[8]: 0.098322319292822702
However, had we not had such a large male sample size, the values might differ greatly; say, for instance, that we had had 3 deletion attempts out of 33 articles on males instead. Then we find In [12]: binom(12, 5/45).sf(1)
Out[12]: 0.39171131326925679
In [13]: fisher_exact([[2, 10], [3, 30]])[1]
Out[13]: 0.59808767522891315
[0]: https://en.wikipedia.org/wiki/Fisher's_exact_testI suppose I could have made a more conservative estimate by not using the female samples in the calculation of 'p'.
I also don't quite like the way Fisher's test fixes the total number of deleted articles, but I guess a more sophisticated test assuming only the total number of female and male articles to be fixed could get a bit complicated.
Deletionists of Wikipedia are just destinied to fall on their face, and embarrass the whole project.
Deletionism is the single big thing that makes Wikipedia somewhat sour in my view.
https://www.meta-chart.com/share/how-wikipedia-deletionism-w...
Two important takeaways are: A lot of stuff on Wikipedia just isn't encyclopedic. For example, there's a lot of sports trivia about obscure teams and local leagues. Such data deserves structured storage and not Wikipedia, and I just don't feel "2013 Richmond Raiders season" is more important than any single matematician who worked with Erdős.
Second takeaway, a whole world of topics is just not ever noted by Wikipedia. For examples, you can have whole branches of local music scenes ineligible for note because every single artist fails to "be mentioned in a paper made of paper" (in 2010s, seriously) or "sell 3k discs made of plastic and foil" (again, this is in 2010s).
If there's some kind of consistent policy I definitely don't see the value of it.
If it's in the New York Times you can probably believe that it's not been created to game Wikipedia, the same cant be said for a random blog.
I regularly see people complain about this without mentioning how they would solve this problem.
I'm not sure what the best answer is either. The closest I've seen is to overlay a more inclusionist Wikipedia with some sort of crowdsourced reputation/quality moderation system. That can be gamed too of course but something along those lines may be the best approach.
I can understand "we could not verify this so we are deleting this article", but not "we speedy deleted this article without checking anything". The latter has no excuses except for obvious commercial spam.
What can go visibly wrong if Wikipedia would cover different subset of obscure topics than it does currently?
We're not talking about faking facts in well-visited articles, since a) that's not what is being discussed, and b) that has been done anyway, getting into news multiple times.
Does deletionism sometimes leads to embarrassing results? As witnessed, it does.
Would less deletionism lead to Wikipedia becoming "not really usable across a wide range of domains"? I can't see how it would, neither could you.
So I would argue that Wikipedia deletion policy is considerably suboptimal.
I guess deletion also serves to keep editors in check, if they don't meet quality standards. It's simply more difficult to delete a deletion-privileged user who has also a number of contributions under the belt. This is similar to police corruption, or shirky's principle.
Also no reason for valid articles to be deleted from Wikipedia.
I am against corruption and this is awful. Nobody should be reasoning this.
It's unlikely most people would do these things just for fun, but allowing citations that can be verified using ordinary research methods seems pretty reasonable.
"We could not verify this fact" is not a valid reason for deletion, let alone speedy deletion. The requirement is verifiability (it must be possible to verify this fact using reliable sources), not verification (it's OK if the sources are not immediately available).
(https://en.wikipedia.org/wiki/Wikipedia:Criteria_for_speedy_...)
They don't just speedy delete articles about music groups, they have a dedicated template for it. As per CSD A7.
I don't see why a page about music group would be "obviously inappropriate". Maybe you will care to explain?
I know, because I've done both.
How about "2013 Richmond Raiders season"? Apparently this is more encyclopedic? There are no specialized sites for this?
I would say you did not carry your point across.
> How about "2013 Richmond Raiders season"? Apparently this is more encyclopedic? There are no specialized sites for this?
No I'm not, I just choose to ignore sportsfandom.
You choose to ignore sportsfandom, but apparently it gets out of jail free card on Wikipedia. Is it a good thing? Why? Where's the motivation? Who decided that?
no it does not. These artist articles are created in subpar condition and left to rot, most of the creators not coming back, in the extreme case but often enough. There's certainly a bunch of obscure artists, say death metal bands ... and on the other hand my local low league team doesn't have a page. The proportions you are comparing are quite different.
> Where's the motivation? Who decided that?
there's a thing or two to say about comercialization, I'll give you that.
And, as you suggest, the fact that the high school quarterback or a local diner got written up in some local newspaper that has managed not to go out of business yet doesn't make them more notable in any important sense than, say, pretty much any scientist with a tenured position at a major university. (Or someone who is well known in open source circles or a senior executive in a major company or ...)
I'm not sure what the best answer is. We pretty much know that anything goes doesn't really work (see Google Knol). But the current system on Wikipedia is pretty flawed as well.
Plenty that really doesn't seem to pass the sniff test manages to stick around. Though I'd prefer too many edge cases than too few. There's already a mechanism to mark pages with a caveat for needing cleanup or expansion after all.
There are certainly plenty of cases that would seem to be at one extreme or the other. National political figures, professional sports teams, etc. etc. are clearly notable. The random person whose name is only visible on the Internet on their own website probably not.
But in between those poles are published but not well-known books, minor league baseball players, little-seen low-budget movies, random board games, amateur bands etc.
Whether, for example, every book that has ever been published could potentially have its own Wikipedia article is ultimately an opinion/philosophy about whether there should be some limits on content that's just very obscure on Wikipedia even if it's theoretically verifiable to at least some degree.
You are not notable just because you had at minimum one collaboration with a famously notable person. She seems to be dissatisfied that she can’t just use the Erdos number as an argument alone to add thin bios for every female collaborator he had. And then goes on to complain that there perhaps exist equally unimportant woman who would not fit into this catagory “just because they didn’t collaborate with Erdos” but just as in-notable as the not notable collaborators...
Stop judgeging notability by Erdos number if that is you singular claim to fame then the correct action is speedy deletion independent of sex.
I don't think I am trolling. There are multiple tools to analyze the statistics, e.g. [1], and the female biographies stood out for me. Everybody checks the deletion discussion stats since a high accuracy (in terms of one's vote matching the consensus) is a de facto requirement to become an admin.
I don't think there is anything sinister going on, simply a few editors with extremely inclusionist views tend to patrol articles about females [2]. Since the average deletion discussion has only ~3 participants, this results in a bias.
[1] https://tools.wmflabs.org/afdstats/
[2] I kid you not, for example Megalibrarygirl, who's recently become an admin, is only involved in articles about women [3]
[3] https://tools.wmflabs.org/afdstats/afdstats.py?name=Megalibr...
What is your definition of "notable"? I assume if you can produce stats you have objective criteria?
Have you deleted "local hero" articles?
What is the criteria for relevance? I recall there being numerous studies in the past demonstrating articles about important events and people in non-western environments being deleted.
Actually, I think there is a significant bias against things that lack sources in English.
Her analysis here seems very shallow and already operating from the preconceived mindset that this must be misogyny. I'm not saying it's not, but the argument is weak.
So the question is, what is the cause for this discrepancy?
It is possible that the women who worked with Erdos were much more likely to be irrelevant than the men -- by a factor of 4 -- in which case this would be reasonable.
The alternative (her claim) is that they were marked for deletion because other editors default to assume that women are less relevant than men, unless proved otherwise.
My question is: are the people who are tagging these articles for deletion field experts -- mathematicians or whatever -- who can actually understand the relative importance of individuals, or are they running through a checklist? If they're just running through a checklist does their checklist just reinforce old stereotypes? For example, old newspapers would not bother interview women engineers, scientists, etc, and even if they were historically women were absolutely kept out of those fields, so even if interviews were random, statistics would ensure low public visibility. Hence relying on the existence of published biographies is not sufficient, and is arguably less relevant than the Erdos number.
Does it treat an Erdos number as being proof of importance in some cases, but not in others? If your default view is men are real mathematicians, and women are not, it is easy to interpret a low Erdos number as evidence of importance (it confirms your belief of relevance), but if you start out with the converse -- this person isn't relevant or important -- then it's just trivia.
Exactly, these are the things you need to analyze before making such a claim. If for example we look at other confounding factors and find that these explain most of the variance, then we can discuss why these other factors might be influenced by gender (e.g. less interest from contemporary news publications).
The statistical point you raise is excellent, though, and deserves more discussion. Besides being male and female, are there other characteristics which distinguish the "notables" from the "non-notables"? Statistical analysis would normally require that the two groups be exchangeable other than being "male" and "female", and we know in advance that this is not the case.
Is Wikipedia to blame if the otherwise valid criteria it chooses for notability (publications, citations) have a disparate impact on women? Probably not, but it's still a valid question of whether a different approach would produce a better societal effect, whatever that might be.
You're probably right. I don't need to bow to PC culture here though. Honesty on her part doesn't preclude her from being biased though. With a topic like this it's very difficult not to be biased. I just wish she had at least assumed her maybe being biased and analyzed the statistics a bit to see whether her intuition stands up to scrutiny. I think it's somewhat probable that it does! But you need to show it.