Did they release any trained model like Google did for word2vec?
Here are the almost-comparable evaluations:
fastText Numberbatch
en:RW .46 .601
en:ws353 .73 .802
fr:rg65 .67 .789
The difference actually should be larger: Numberbatch considers missing vocabulary to be a problem, and takes a loss of accuracy accordingly, while FastText just dropped their out-of-vocabulary words and reported them as a separate statistic.I'm using their Table 3 here. I don't know how Table 2 relates, or why their French score goes down with more data in that table.
What's the trick? Prior knowledge, and not expecting one neural net to learn everything. Numberbatch knows a lot of things about a lot of words because of ConceptNet, it knows which words are forms of the same word because it uses a lemmatizer, and it uses distributional information from word2vec and GloVe.