quoted performance numbers on multiple GPUs leave other frameworks in the dust. where's the catch?
No idea what the catch is :)
If so, while very cool, that's not a general solution. Scaling batch sizes of 256 or lower would be the breakthrough. I suspect they get away with this because speech recognition has very sparse output targets (words/phonemes).
Too bad the code below isn't open-source because they got g2 instances with ~2.5 Gb/s interconnect to scale:
http://www.nikkostrom.com/publications/interspeech2015/strom...
Training data throughput isn't the right metric to compare -- look at time to convergence, or e.g. time to some target accuracy level on held-out data.