Why are GPUs well-suited for deep learning? (2016)
quora.com
quora.com
The other interesting property of 3D rendering is that it is an 'embarrassingly parallel' problem. Vertices or pixels (mostly) don't depend on each other, so the GPU can process them in parallel. That's how GPUs turned over time into big arrays of very simple cores which are good at running a huge number of vector operations in parallel.
1. Matrix Multiplication caches easily: you can program the data-fetches to occur in blocks and recycle your memory-fetches across a huge number of operations.
2. Matrix Multiplication is innately parallel, with every square of the output matrix possibly computed in parallel.
-----------
That being said: GPUs are not the best at this. Systolic processors are the best matrix multipliers and that's where NVidia Tensor Cores and/or Google TPUs come in.
------------
GPU / SIMD architecture was chosen for GPUs because video games have a large number of 4x4 matrix multiplications per frame. Every single vertex, of which a typical video game could have ~100 million vertexes, needs to be 4x4 matrix multiplied to determine where on the screen it has moved.
GPUs need to perform this operation at 60 FPS for video gamers to be pleased, though some video gamers demand 120 FPS or even 200+ FPS.
4x4 is the operation of an affine transformation in 3d space (translation, rotation, scaling, etc. etc.). Most commonly for 'camera movement'. (Player moved to the right 5 ticks and then looked to the north-west by 27 degrees, where should all the objects in the screen move to? This is solved by matrix-multiplying all 100-million verticies with the camera-position matrix)
4x4 Matrix Multiplication isn't quite big enough to "deserve" a systolic processor yet.
"As many have said GPUs are so fast because they are so efficient for matrix multiplication and convolution, but nobody gave a real explanation for why this is so. The real reason for this is memory bandwidth and not necessarily parallelism."
Deep Learning chose TPUs / Tensors because GPUs (and other computers) are faster at matrix multiplication than any other operation. Many many other representations of neural networks have existed over the decades, its the matrix-based "Tensor" that performed the best and took off.
If you ask "why are GPUs better suited than CPU s?" then it's because CPUs are more general purpose, and need to handle branching, task-switching, IO.
For rendering graphics, GPUs need to apply the same set of instructions to multiple pixels and vectors, so they use a "single instruction, multiple threads" architecture. They apply the same operation to different data, which allows them to do this in parallel, so they can have a lot of ALUs.
And this SIMT architecture is also really good for matrix multiplication. Many of the features that make CPUs fast - caching, pipelining - wouldn't work with SIMT.
The important tidbit is that CPUs find parallelism where the programmers didn't specify any. CPUs are "out of order" machines, spending a fair amount of energy and internal storage (the ROB / reorder buffer) to execute out-of-order, and then place all the results back in order before the programmer notices.
GPUs, instead of being configured to find parallelism, innately assume that the programmer laid out parallelism for it ahead of time. GPUs can only work with a massively parallel language (OpenCL, DirectX HLSL, or CUDA) specifying thousands of parallel work units per line.
Gracious. Put a stamp on posts, Quora!