Although it will give you more or less reasonable results, I don't see a null hypothesis for which this approach does not ultimately become flawed: Assuming that you are testing the hypothesis that the two proportions agree, then, for instance, using a fixed binomial parameter of 12/262 becomes wrong for any number of deletion attempts different from 2; that is, part of the assumed model includes the outcomes you are testing: The case "3 deletions", which is part of the "2 or more extreme" probability you calculate, would have led to a different value of $p$.
Various statistical tests exist for taking this into account. With Fisher's exact test [0], we find a $p$-value close to yours:
In [6]: from scipy.stats import fisher_exact, binom
In [7]: binom(12, 12/262).sf(1)
Out[7]: 0.10211478280975775
In [8]: fisher_exact([[2, 10], [10, 240]])[1]
Out[8]: 0.098322319292822702
However, had we not had such a large male sample size, the values might differ greatly; say, for instance, that we had had 3 deletion attempts out of 33 articles on males instead. Then we find In [12]: binom(12, 5/45).sf(1)
Out[12]: 0.39171131326925679
In [13]: fisher_exact([[2, 10], [3, 30]])[1]
Out[13]: 0.59808767522891315
[0]: https://en.wikipedia.org/wiki/Fisher's_exact_test