Why GEMM is at the heart of deep learning
petewarden.com
petewarden.com
[1]: http://kukuruku.co/hub/algorithms/using-the-quick-raise-of-m... [2]: http://fortranwiki.org/fortran/show/Fortran+2008
2. You can call these libraries from any language, so Fortran offers no advantage there.
3. Fortran performance is similar to C and C++ (assuming use of restrict where needed).
4. CUDA will not lead to 10-40x boost unless your CPU implementation is garbage. The realistic standard is about 2x when normalized by Watts; see the Green500 for example. The 2x also holds for price when using server-grade parts (like Tesla), but GPUs come out further ahead if you buy consumer grade. Note that you still have to pay to buy and power a host, and that GPUs are a lot less versatile in terms of problem size/turn-around time and workload. They have their place, but evaluate your workload carefully before concluding that the cost proposition favors GPUs.
2. If I were to run a highly iterative algorithm, say with 100,000 or more iterations, would the memory marshaling not be significant? Or would it be unnoticeable since I would be using these libraries anyway? I also personally find the array notation and slicing to be quite nice.
3. The use of "restrict" is nice. I was not aware of that feature. Looking at this reference [1], it seems that Fortran basically does the equivalent of automatically using "restrict".
4. According to NVIDIA [2], cuBLAS outperforms MKL by 6x-17x (it seems MKL is catching up since I last checked). Is this misleading by NVIDIA? What is your expectation for a reasonable execution speed increase? The fixed-cost efficiency and variable-cost efficiency are good points to bring into the discussion. Your point about carefully evaluating workload is also a good one. I don't want people to think that throwing an algorithm into CUDA will magically speed up the overall program. Network I/O, disk I/O, and device (GPU) I/O are all important bottlenecks to consider.
3. Yes.
4. NVIDIA has a track record of misleading comparisons. They cleaned up their act in the these CUDA-7 benchmarks: http://devblogs.nvidia.com/parallelforall/cuda-7-release-can... in response to this G+ discussion https://plus.google.com/+JeffHammondScience/posts/G1MzHqZaxy... . Thanks to Szilárd Páll for calling attention to this discussion and to Mark Harris for updating the plots. Note that this should not be interpreted as "MKL/Xeon is catching up" but rather that performance comparisons are sensitive to the details of the experiment and it can be hard to recognize the consequences of biased configurations. Better normalization and standards for comparison can help.
Also, while the original Blas was in Fortran, you probably want OpenBlas, which is in C. Or maybe Intels MKL.
And since they claim fbfft was even faster than cufft, the situation should look even better for fft.
EDIT: And they are about tied in the 3x3, 64 case.
"The main competitor to the GEMM approach is using Fourier transforms to do the operation in frequency space, but the use of strides in our convolutions makes it hard to be as efficient."
Kind of an embarrassing typo for an article that is essentially about matrix multiplication.
Here is a corrected (transposed) version of that figure: http://i.imgur.com/Bo6CJCv.png
Edit: It looks like every figure in the article suffers from this problem. I know FORTRAN matrices store each column contiguously in memory, but whether a matrix is accessed like * (ptr + r * ncols + c) or * (ptr + c * nrows + r) is (in my opinion) a low-level detail and shouldn't affect our thinking about how multiplication works.
EDIT: yep, the figures were revised. Compare the corrected version [1] vs the original [2].
[1] https://petewarden.files.wordpress.com/2015/04/gemm_correcte...
[2] https://petewarden.files.wordpress.com/2015/04/gemm.png
EDIT 2: I completely missed that the author put a notice (about having revised the figures) at the bottom of the post.
Not sure if this holds true for other NN structures like GoogLeNet or the weirder recent fully convolutional networks.
in general, to make deep learning faster avoid using backpropagation
It's important to note that (from what I understand) this paper seems to focus on building single-layer NNs, which are quite different from the behemoth AlexNet that this article focuses on.
For very large parameter spaces, gradient based optimization is really the only game in town. There's also lots of clever tricks (gradient clipping, momentum, adagrad, etc) that are involved in optimizing the convergence.
As for making training faster, second-order methods don't generally work as well as backpropagation for large networks. Some notes from Yann LeCunn! http://yann.lecun.com/exdb/publis/pdf/lecun-98b.pdf