Short answer is actually: maybe!
The bulk of the work done in this code (in terms of FLOPS and, likely, wall-clock time) is going to be in BLAS-3 operations in the feed-forward and back-prop steps. That is, almost all of the work is done using Matrix-Matrix multiplies and in-place arithmetic/transcendental functions.
CUBLAS[1] will allow you to run these types of operations on your GPU at highly accelerated rates, without much more effort than replacing your BLAS library with a new binary. Additionally, if you want finer granularity control over what gets done on the GPU, there are other libraries[2] which provides a direct interface to CUBLAS.
[1] https://developer.nvidia.com/cublas [2] https://github.com/JuliaGPU/CUBLAS.jl