I remember writing the RBM in straight CUDA back in ~2008. Back then we had to do our tensor work uphill both ways against the gradient.
ML researchers following trends, like doing end-to-end CNN for whatever problems, and using this LSTM and that GRU and that latest architectures; in similar fashion like web devs picking and dropping javascript frameworks.
Most "off the shelf" work people do is fairly simple transfer learning on an imagenet cnn.
That or if they do generative, most of the hype is focused on GANs now.