> Our approach requires an expensive pre-training step - 1 month on 8 GPUs.
Wow! I enjoy playing with neural networks but this kind of thing reminds me that I'm not really doing deep learning...
I have no idea how researchers could have the patience and confidence to wait that long for a result. In my own (small-data) work, I get frustrated if it doesn't converge in half an hour.. I constantly end up Ctrl-C'ing and tweaking things if it doesn't behave as expected, or appear to be continuing to improve.