Other researchers have noted[2] that with a set of command line flags that vw is almost the same as the system described in the paper, specifically, "vw --ngrams 2 --log_multi [K] --nn 10".
Behind the speed of both methods is use of ngrams^, the feature hashing trick (think Bloom filter except for features) that has been the basis of VW since it began, hierarchical softmax (think finding an item in O(log n) using a balanced binary tree instead of an O(n) array traversal) and using a shallow instead of deep model.
I am still interested in the more detailed insights the team from Facebook AI Research may provide but the initial paper is a little light and they're still in the process of releasing the source code.
^ Illustrating ngrams: "the cat sat on the mat" => "the cat", "cat sat", "sat on", "on the", "the mat" - you lose complex positional and ordering information but for many text classification tasks that's fine.
[1]: https://github.com/JohnLangford/vowpal_wabbit/wiki
[2]: https://twitter.com/haldaume3/status/751208719145328640