Fun fact: Long before the dawn of the GPU deep learning hype (and even before CUDA was a thing), a bunch of CS nerds from Korea managed to train a neural network on an ATI (now AMD) Radeon 9700 Pro using nothing but shaders [1]. They saw an even bigger performance improvement than Hinton and his group did for AlexNet 8 years later using CUDA.
[1] https://ui.adsabs.harvard.edu/abs/2004PatRe..37.1311O/abstra...