I strongly recommend fast.ai instead. Although often looked at as the resource for people who can’t deal with the math, I actually found it to be extremely good at explaining the math. Compare, for example, the deep learning book’s explanations on various gradient descent methods with Jemery Howard’s explanation - in the book it looks very complex, whereas in the course it’s actually really intuitive. And Jeremy doesn’t gloss over things, he actually implements the various gradient descent methods in Excel (!).
I started with a top-down approach via the fast.ai courses and learning Keras, then spent time brushing up on some of the math concepts (as you said, it assumes a fair amount of previous knowledge), and then went back and started re-reading it, and I finally feel like I'm starting to get some value out of it.
Definitely wouldn't recommend it as a first book, though.
That book is only useable if you are a math PhD and want to get into ML.
Source: buddy is math PhD and worked with ML for 5-10 years now, even he has hard time understanding some chapters.
I did find that it didn't provide much context around why the equations matter, and definitely wouldn't be useful for those starting out in the field. It did have some pretty good coverage of gradient descent and various optimisers, which I found useful.
tl;dr: not really worth it for its stated purpose, but not a bad second or third stats book.