This is how i'd describe it. Deep learning is a set of tinker toys. Lego blocks if you will that you can sculpt with data into some very interesting models. Its an art, where the brushstrokes are matrices. Place an attention module here, and a convolution net there. And throw in a tensor with a softmax, and viola.
Now I love math. and part of me really wants to see deep learning become a mathematical discipline. Deep in the backwaters there are parts of deep learning involve some math (think variational inference, bayesian models, etc). And I do want deep learning to be about condition numbers and combinatorics. But if you want to be perfectly honest with a newbie in the field, if you want to get your feet wet in deep learning, don't waste 3 months on a class on advanced optimization or measure theory or probability. Just dive in
The only reason to fret about this in my opinion is if you're a PhD machine learning engineer who doesn't want the field to open up to non-PhDs. I think data scientists are used to being able to say, "hey, if you don't have a PhD, you really can't do or understand what I do" -- deep learning represents potentially a huge culture shock to that attitude. But even the Google Brain research team has some non-PhDs now.
I do think deep learning practitioners should learn the math. I just don't think there's actually that much math to learn. Certainly, if you read through the TensorFlow MNIST tutorial and you have no idea what cross-entropy is and you don't understand what the softmax layer is for, you need to go back to the basics. But these are concepts that anyone with any reasonable engineering degree can pick up relatively quickly.
As an example, I submit the five articles on Distill, a new online machine learning journal:
Notice that only the first has any real math, and even there the math is just not very advanced -- it's undergraduate-level.
I have dabbled in writing a super simple neural network to solve the MNIST. Using an example written in python and porting it to go, so that I couldn't copy and paste. I had to see what each step did. It very rapidly went to about 30% accuracy and stuck there. So I know I did something wrong. But I abandoned it after not being able to figure out what.
It's definitely the best reference on the subject. With only calculus 3 under your belt the math won't be trivial, but it should overall be fairly approachable and certainly much more so than something like "The Elements of Statistical Learning".
If a student wants to learn how to play the guitar, you show them 3 chords so they can play Bob Marley or Oasis.
You don't require them to first study consonance, dissonance, rhythm, melody, timbre, dynamics, articulation, texture, form, expression, notation, song writing, Schenkerian analysis, harmonic identity, semiotics, and musical set theory.
Someone who can play the guitar with a passion, can be taught to learn musical notation. The other way around is not guaranteed.
Your suggestion is not necessarily bad: It's good to learn the maths about the Wasserstein metric if you are using GAN's. But for effective teaching your suggestion is archaic, and part of the mindset that makes student's eyes glaze over when being taught mathematics. Can you point to a success story of a student to neural network researcher that did not start with a practical application?
Hyperparameter optimisation is basically a fudge right now - you try everything and see what works. Even the research groups who came up with the standard network stacks, like VGG, basically lucked out and found an architecture that worked, then tried several variants and found one that worked better. DL papers are full of handwaving speculation about why particular networks perform better than others, but right now it's just that: highly educated speculation.
This isn't limited to deep learning. If you want to try any kind of machine learning, it's totally reasonable to throw different fitting functions at your problem to see which one works best. Unless you have an unusually clear problem category, it's rarely possible to say at the outset that "This problem would best be solved with method <X>". A counter here would be that if you need to classify images, you should almost certainly use a convnet.
You need some understanding about why things might be going wrong, e.g. your loss isn't moving -> crank up the learning rate. You're seeing nans? Probably your learning rate is too high. But that doesn't really need any serious maths to understand. You can get by quite well by figuring out empirical rules.
I'm not arguing that you shouldn't learn the maths, it's a wise idea to, but many people use deep learning models without knowing how backpropagation works for instance.