[1]: http://www.wolframalpha.com/input/?i=1+-+(1+-+12%2F262)%5E12...
[1]: http://www.wolframalpha.com/input/?i=1+-+(1+-+12%2F262)%5E12...
Only 17.47% (as of 26 February 2018) of the English Wikipedia's biographies are about women: https://en.wikipedia.org/wiki/Wikipedia:WikiProject_Women_in...
EDIT: I rechecked the quote from the article about the likelihood of deletion and found it to be "in this cohort of Erdős Ones, the pages for women were over four times more likely than men’s to be nominated for speedy deletion." (my italics). FYI.
Wikipedia being a tertiary source may, through unbiased processes, simply be reflecting bias in the secondary sources. Biographies of pre-modern women are far less common than of their male contemporaries, IME; that's less true of my modern and especially extremely current figures, but Wikipedia isn't exclusively focused on current figures.
However if it were 17.47% for current times, that would be a significant bias that needs to be addressed. It still might not hit 50/50 given reduced opportunities for women still today, but it should be moving ever closer.
Talking about p-values in contexts where one has all the facts and the additional fact that one has all the facts suggests a gross misunderstanding of statistical concepts, their utilities, and the circumstances in which they apply.
First of all I at least don't have all the information. If I had it would indeed be meaningless to infer the reason behind the deletion attempts, as I'd know already, but the fact remains that I don't. By this logic trying to see how likely this was to happen assuming articles were just randomly marked for deletion is a perfectly reasonable test.
Second of all, how exactly is this situation different from tossing a coin several times? If I toss a coin 50 times, record the results, and then melt down the coin, surely it'd still be valid to conclude that it was biased if it came up heads 49 times out of 50? Or conversely that it wasn't biased if it came up tails 26 times out of 50?
Arguing that statistics only applies to experiments that can be repeated suggests a gross misunderstanding of statistical concepts, their utility and the circumstances in which they apply.
Maybe I'm misunderstanding Ben's point, but I don't think he's objecting to this as a frequentist. You have correctly calculated that given the small sample size of female candidates, there is about a 10% chance that random deletion of entries would produce a disparity this large.
I think Ben's issue is that since we already know that the deletion process is nonrandom, this tells us nothing we do not already know. The question we're actually interested in is what the factors are that cause deletion requests, not just the ones that correlate with them, and the numerical value you've calculated doesn't give us much insight.
So maybe you could talk more about what you think the calculated p-value actually means, and what we can do with it?
Viewed from that perspective what this p-value means is that merely knowing the proportion of articles that gets marked for deletion is already enough to conclude that it is likely (or at least not impossible) for a group of 21 articles to have at least 2 articles marked for deletion. In that sense the knowledge that of the 21 articles written on women, 2 were marked for deletion doesn't really tell us something that we didn't already know.
Of course that doesn't mean we've exhausted all available knowledge, but you're going to have to look further than just how many articles get deleted.
Various statistical tests exist for taking this into account. With Fisher's exact test [0], we find a $p$-value close to yours:
In [6]: from scipy.stats import fisher_exact, binom
In [7]: binom(12, 12/262).sf(1)
Out[7]: 0.10211478280975775
In [8]: fisher_exact([[2, 10], [10, 240]])[1]
Out[8]: 0.098322319292822702
However, had we not had such a large male sample size, the values might differ greatly; say, for instance, that we had had 3 deletion attempts out of 33 articles on males instead. Then we find In [12]: binom(12, 5/45).sf(1)
Out[12]: 0.39171131326925679
In [13]: fisher_exact([[2, 10], [3, 30]])[1]
Out[13]: 0.59808767522891315
[0]: https://en.wikipedia.org/wiki/Fisher's_exact_testI suppose I could have made a more conservative estimate by not using the female samples in the calculation of 'p'.
I also don't quite like the way Fisher's test fixes the total number of deleted articles, but I guess a more sophisticated test assuming only the total number of female and male articles to be fixed could get a bit complicated.