The dataset
is biased. Otherwise woman~nurse wouldn't have been the second hit for the query. Had the dataset
not been biased it would have been able to produce the analogy "woman is to doctor as man is to nurse". See table 4b where they perform that experiment.
I think it is also important to note that this is an argument over which of two methodologies is best. The answer likely is that it depends on the use case. Researchers are not being accused of using underhanded methods, manipulating data or cheating.