Open source deep learning models that programmers can download and run first try
github.com
github.com
Besides this point, there is much more than simple derivatives in deep learning. For example regularization can yield quadratic programming problems. Different optimization algorithms can have tremendous impact on training time and model performance. Models can be quite sensitive to specific parameters that you can't just set at random.
More ingenious architectures like GAN also require some fairly technical thinking to get right. There is much more than image classification and vanilla NN or CNNs.
But then why does using 5 layers work worse than 4? Your theory is no good at predicting what the hyperparameters should be. The only way to find the correct hyperparameters is through empirical search.
>there is much more than simple derivatives in deep learning. For example regularization can yield quadratic programming problems. Different optimization algorithms can have tremendous impact on training time and model performance.
All these concepts are fairly simple also and can be expressed with little math. Additionally, a casual user doesn't need to have a deep understanding of them and the library will usually take care of it. Any more than a programmer needs to have a deep understanding of how an optimizing compiler works.
>More ingenious architectures like GAN also require some fairly technical thinking to get right.
The idea of using NNs to trick each other, is also fairly simple. It doesn't even involve any math.
The other thing is the unfortunate/misleading/atrocious jargon that has been adopted.
Mathematical notation is basically a programming language. A programming language with weird symbols you can't type to search for, single letter variable names for everything, and no comments. And it's written by programmers that are obsessed with fitting everything into a simple line and making it as small as possible, no matter how difficult it is to read. Any programmer understands this is incredibly bad practice. And even if parse every step and perfectly follow what the code is doing, without explanation, it's pretty difficult to figure out why.
A very bad one that can only be executed by brains with the requisite existing historical knowledge; in fact it's more like bad pseudo-code that lacks the explicitness necessary to translate into actual instructions. It's basically condensed jargon intended for the already converted.
It'd probably be vastly easier to teach math with an actual programming language than with traditional notation. Scheme would be ideal for this.
It's about the level of abstraction. And yeah if you don't understand the notation or syntax at the level of abstraction you're studying, it will be very hard.
(FWIW I find Scala code quite hard to understand sometimes, but I also find the more I know about the language, the more comprehensible it gets).
It doesn't matter how familiar you are with the language. Without an explanation of what the hell is going on, just looking at the code is useless.
That said, it can be useful for the beginner to implement a basic NN library from the ground up, to understand how the vector processing works, as well as what is going on step-wise with backprop and such.
Once that is fully understood, the next step of utilizing true vector processing libraries can be taken, and so on - eventually culminating in using and understanding libraries like TensorFlow.
Having the background of the lower levels gives you an appreciation and even insights when you transition to higher level frameworks.
That's just my opinion, though.
I definitely agree though that it's more of an experimental science at the moment.
Mathematically, closest to that would be Hilbert's program.
Though neural nets can paint like Van Gogh nowadays, asking them to come up with Hilbert's program may be a bit too much of an ask. Yet I would not deeply mind if researchers would revisit papers like http://www.ics.uci.edu/~rickl/publications/1996-icml.pdf "On the Learnability of the Uncomputable".
Moreover, I would probably encourage people to read examples of Tensorflow or Caffe2 running on iOS rather than something like Forge. Forge is an interesting project but won't really help you if you don't have a clue about MPS or Deep Learning.
This is how i'd describe it. Deep learning is a set of tinker toys. Lego blocks if you will that you can sculpt with data into some very interesting models. Its an art, where the brushstrokes are matrices. Place an attention module here, and a convolution net there. And throw in a tensor with a softmax, and viola.
Now I love math. and part of me really wants to see deep learning become a mathematical discipline. Deep in the backwaters there are parts of deep learning involve some math (think variational inference, bayesian models, etc). And I do want deep learning to be about condition numbers and combinatorics. But if you want to be perfectly honest with a newbie in the field, if you want to get your feet wet in deep learning, don't waste 3 months on a class on advanced optimization or measure theory or probability. Just dive in
The only reason to fret about this in my opinion is if you're a PhD machine learning engineer who doesn't want the field to open up to non-PhDs. I think data scientists are used to being able to say, "hey, if you don't have a PhD, you really can't do or understand what I do" -- deep learning represents potentially a huge culture shock to that attitude. But even the Google Brain research team has some non-PhDs now.
I do think deep learning practitioners should learn the math. I just don't think there's actually that much math to learn. Certainly, if you read through the TensorFlow MNIST tutorial and you have no idea what cross-entropy is and you don't understand what the softmax layer is for, you need to go back to the basics. But these are concepts that anyone with any reasonable engineering degree can pick up relatively quickly.
As an example, I submit the five articles on Distill, a new online machine learning journal:
Notice that only the first has any real math, and even there the math is just not very advanced -- it's undergraduate-level.
I have dabbled in writing a super simple neural network to solve the MNIST. Using an example written in python and porting it to go, so that I couldn't copy and paste. I had to see what each step did. It very rapidly went to about 30% accuracy and stuck there. So I know I did something wrong. But I abandoned it after not being able to figure out what.
It's definitely the best reference on the subject. With only calculus 3 under your belt the math won't be trivial, but it should overall be fairly approachable and certainly much more so than something like "The Elements of Statistical Learning".
If a student wants to learn how to play the guitar, you show them 3 chords so they can play Bob Marley or Oasis.
You don't require them to first study consonance, dissonance, rhythm, melody, timbre, dynamics, articulation, texture, form, expression, notation, song writing, Schenkerian analysis, harmonic identity, semiotics, and musical set theory.
Someone who can play the guitar with a passion, can be taught to learn musical notation. The other way around is not guaranteed.
Your suggestion is not necessarily bad: It's good to learn the maths about the Wasserstein metric if you are using GAN's. But for effective teaching your suggestion is archaic, and part of the mindset that makes student's eyes glaze over when being taught mathematics. Can you point to a success story of a student to neural network researcher that did not start with a practical application?
Hyperparameter optimisation is basically a fudge right now - you try everything and see what works. Even the research groups who came up with the standard network stacks, like VGG, basically lucked out and found an architecture that worked, then tried several variants and found one that worked better. DL papers are full of handwaving speculation about why particular networks perform better than others, but right now it's just that: highly educated speculation.
This isn't limited to deep learning. If you want to try any kind of machine learning, it's totally reasonable to throw different fitting functions at your problem to see which one works best. Unless you have an unusually clear problem category, it's rarely possible to say at the outset that "This problem would best be solved with method <X>". A counter here would be that if you need to classify images, you should almost certainly use a convnet.
You need some understanding about why things might be going wrong, e.g. your loss isn't moving -> crank up the learning rate. You're seeing nans? Probably your learning rate is too high. But that doesn't really need any serious maths to understand. You can get by quite well by figuring out empirical rules.
I'm not arguing that you shouldn't learn the maths, it's a wise idea to, but many people use deep learning models without knowing how backpropagation works for instance.
Half of this stuff I can't run.