To say that there were troublesome concerns about the evaluation is a bit too strong. Yoav's experiment showed that for the one data setup he ran, training both models on the same data, GloVe out-performed word2vec by less than in our results where we compared against the publicly released word2vec vectors. But it still outperformed word2vec. And Yoav's comparison isn't the last word: In his experiments, he ran GloVe for only 15 iterations, but, as we already knew and were taking advantage of, GloVe's performance continues to improve for many more iterations. This is now much more clearly documented in the final version of the paper (see fig. 4).
But at the end of the day, these numeric differences were never the point of the paper. The contribution of the paper is to show how the kind of good results word2vec gets with online learning on a token stream can be achieved also by working from a global co-occurrence count matrix, more in the style of the traditional SVD, but changing the loss function and frequency scaling, and that you could expect working in this way to be somewhat more statistically efficient. Yoav has actually been involved in some very interesting work along the same lines himself: https://levyomer.files.wordpress.com/2014/09/neural-word-emb...
@jo_9: We don't use word2vec for training, only for experimental comparison, and the style of training is fairly different.