I don't think it's usable IRL (in its present state), as, according to figures, it doesn't work well even on cifar/mnist. Correct me if I'm wrong, but the value of this paper is that you can decouple a model and train layers asynchronously/independently, just first steps to distributed NN training.
Well, HogWild works well enough. It's another one for the giant "weird shit you can get away with in neural network land" bucket, I think
in the paper, it seemed to work quite well for LSTMS (language modeling), although it is best suited for tasks with very, very large neural networks for which parallelism is beneficial