"As many have said GPUs are so fast because they are so efficient for matrix multiplication and convolution, but nobody gave a real explanation for why this is so. The real reason for this is memory bandwidth and not necessarily parallelism."
Deep Learning chose TPUs / Tensors because GPUs (and other computers) are faster at matrix multiplication than any other operation. Many many other representations of neural networks have existed over the decades, its the matrix-based "Tensor" that performed the best and took off.
The other interesting property of 3D rendering is that it is an 'embarrassingly parallel' problem. Vertices or pixels (mostly) don't depend on each other, so the GPU can process them in parallel. That's how GPUs turned over time into big arrays of very simple cores which are good at running a huge number of vector operations in parallel.
1. Matrix Multiplication caches easily: you can program the data-fetches to occur in blocks and recycle your memory-fetches across a huge number of operations.
2. Matrix Multiplication is innately parallel, with every square of the output matrix possibly computed in parallel.
-----------
That being said: GPUs are not the best at this. Systolic processors are the best matrix multipliers and that's where NVidia Tensor Cores and/or Google TPUs come in.
------------
GPU / SIMD architecture was chosen for GPUs because video games have a large number of 4x4 matrix multiplications per frame. Every single vertex, of which a typical video game could have ~100 million vertexes, needs to be 4x4 matrix multiplied to determine where on the screen it has moved.
GPUs need to perform this operation at 60 FPS for video gamers to be pleased, though some video gamers demand 120 FPS or even 200+ FPS.
4x4 is the operation of an affine transformation in 3d space (translation, rotation, scaling, etc. etc.). Most commonly for 'camera movement'. (Player moved to the right 5 ticks and then looked to the north-west by 27 degrees, where should all the objects in the screen move to? This is solved by matrix-multiplying all 100-million verticies with the camera-position matrix)
4x4 Matrix Multiplication isn't quite big enough to "deserve" a systolic processor yet.