Not if there is a difference in outcome, based on gender; if based on the same background, people from gender A publish on average four papers and people from gender B publish on average three papers during their degree, which gender does it make sense to pick? Or if you add a person from gender C to a group that was composed solely of gender D, does it increase or decrease the lab's total output because of different social dynamics? Or, if there are grants that are only accessible to gender E due to positive discrimination, does it make sense to lower their pay, since they can compensate through that grant?
The paper assumes independence between the gender of the applicant and the gender of the students of the professor, which isn't the case unless you hide the applicant in a room away from the current grad students.
Furthermore, the paper mentions that they used both Student's t-test and an ANOVA, both of which assume that the underlying data is normally distributed; there was no mention of a normality test anywhere in the paper. If the data is not normally distributed, which could be the case, then it's violating the normality assumption of both tests, which could potentially make them invalid.
Finally, are there other confounding factors in the study? For example, if the applications were sent in different batches, the time at which they arrived does have an influence. For example, a recent paper evaluated the decision of judges for parole and found out that judges were more lenient right after lunch and were much more strict right before[1].
As they say, further study is required.
[1] http://www.pnas.org/content/108/17/6889