Besides this point, there is much more than simple derivatives in deep learning. For example regularization can yield quadratic programming problems. Different optimization algorithms can have tremendous impact on training time and model performance. Models can be quite sensitive to specific parameters that you can't just set at random.
More ingenious architectures like GAN also require some fairly technical thinking to get right. There is much more than image classification and vanilla NN or CNNs.
But then why does using 5 layers work worse than 4? Your theory is no good at predicting what the hyperparameters should be. The only way to find the correct hyperparameters is through empirical search.
>there is much more than simple derivatives in deep learning. For example regularization can yield quadratic programming problems. Different optimization algorithms can have tremendous impact on training time and model performance.
All these concepts are fairly simple also and can be expressed with little math. Additionally, a casual user doesn't need to have a deep understanding of them and the library will usually take care of it. Any more than a programmer needs to have a deep understanding of how an optimizing compiler works.
>More ingenious architectures like GAN also require some fairly technical thinking to get right.
The idea of using NNs to trick each other, is also fairly simple. It doesn't even involve any math.
The other thing is the unfortunate/misleading/atrocious jargon that has been adopted.
Mathematical notation is basically a programming language. A programming language with weird symbols you can't type to search for, single letter variable names for everything, and no comments. And it's written by programmers that are obsessed with fitting everything into a simple line and making it as small as possible, no matter how difficult it is to read. Any programmer understands this is incredibly bad practice. And even if parse every step and perfectly follow what the code is doing, without explanation, it's pretty difficult to figure out why.
A very bad one that can only be executed by brains with the requisite existing historical knowledge; in fact it's more like bad pseudo-code that lacks the explicitness necessary to translate into actual instructions. It's basically condensed jargon intended for the already converted.
It'd probably be vastly easier to teach math with an actual programming language than with traditional notation. Scheme would be ideal for this.
It's about the level of abstraction. And yeah if you don't understand the notation or syntax at the level of abstraction you're studying, it will be very hard.
(FWIW I find Scala code quite hard to understand sometimes, but I also find the more I know about the language, the more comprehensible it gets).
It doesn't matter how familiar you are with the language. Without an explanation of what the hell is going on, just looking at the code is useless.
That said, it can be useful for the beginner to implement a basic NN library from the ground up, to understand how the vector processing works, as well as what is going on step-wise with backprop and such.
Once that is fully understood, the next step of utilizing true vector processing libraries can be taken, and so on - eventually culminating in using and understanding libraries like TensorFlow.
Having the background of the lower levels gives you an appreciation and even insights when you transition to higher level frameworks.
That's just my opinion, though.
I definitely agree though that it's more of an experimental science at the moment.
Mathematically, closest to that would be Hilbert's program.
Though neural nets can paint like Van Gogh nowadays, asking them to come up with Hilbert's program may be a bit too much of an ask. Yet I would not deeply mind if researchers would revisit papers like http://www.ics.uci.edu/~rickl/publications/1996-icml.pdf "On the Learnability of the Uncomputable".