Matrix Calculus for Deep Learning
parrt.cs.usfca.edu
parrt.cs.usfca.edu
But more importantly - I need to mention that Terence Parr did nearly all the work on this. He shared my passion for making something that anyone could read on any device to such an extent that he ended up creating a new tool for generating fast, mobile-friendly math-heavy texts: https://github.com/parrt/bookish . (We tried Katex, Mathjax, and pretty much everything else but nothing rendered everything properly).
I've never found anything that introduces the necessary matrix calculus for deep learning clearly, correctly, and accessibly - so I'm happy that this now exists.
The only tensor notation I've been happy with is that used in J (http://www.jsoftware.com), which is simple, flexible, and concise.
There's also some nice-enough modern notation used in this excellent review: http://www.cs.cmu.edu/~christos/courses/826-resources/PAPERS...
A_ij = B_ik C_kj
Differentiating this with respect to variable l: ∂_l A_ij = ∂_l (B_ik C_kj) = (∂_l B_ik) C_kj + B_ik (∂_l C_kj)
By writing out indices you can just use the rules for scalar derivatives.If we're talking about Einstein notation, then I'm a fan - `np.einsum()` is often a great way to create fast tensor computations with minimal code.
As an extra minor nit, italicizing functions like sin, etc. is also somewhat unconventional in mathematical typesetting.
@media (max-width: 768px) {
p {
font-size: 1rem;
}
}Eg. In the first line in 'References'
HTML: "When looking for resources on the web, search for “matrix calculus” not “vector calculus.” "
Latex/pdf: "When looking for resources on the web, search for “elements” not “elements” "
Mathematica doesn't seem to be able to do matrix calculus, which surprised me quite a bit.
f[x_,y_] := x^2 + Sin[y]
vars={x,y};
Table[ D[f[x,y],var1, var2], {var1,vars}, {var2,vars}] // MatrixForm
A = {{1,2},{3,4}}
vec = {x^2, x^3}
D[vec.A.vec, x]
Or perhaps like this, again the table of derivatives:
xvec = {x1,x2}
Table[ D[xvec.A.xvec,x] ,{x,xvec}]
(all untested... one typo caught...)
aa[x_] = {{1, 2}, {3, 4}} x
bb[x_] = {x^2, x^3}
D[ a[x].b[x] , x, x] (* for any suitable tensors *)
% /. {a -> aa, b -> bb}
The math is super easy but keeping all the notation s and conventions in my head is hard, I've never seen it laid out this nicely before. Thanks!
Index notation also seems natural for programming: an element A[i,j] or a slice Z[3,4,:] are precisely this.
That’s... not that sexy. But at least it makes sense to anyone with an undergrad degree in CS or math, which is something neural networks never accomplished.
Are you pulling my leg here, or do I need to scrap my understanding of calculus?
If they exist, how could they not be equivalent?
If anyone can recommend any books, courses, or any other material that starts from high-school level math, and gradually increases in complexity, I would love to look at it.
Cheers :)
Fair warning, I have shown it to a programmer who claimed some level of "math phobia," and they said the first chapter was too difficult. I rewrote that chapter since, and I think it is better, but I could use some feedback :)
If you give it a go and find you're not successful, I'd be interested to hear where is the first point where you got stuck and couldn't get unstuck, since that would suggest a need for us to improve our paper!
I greatly appreciate your response! I will take a long look at this paper again and attempt to digest it.
Thanks again!
https://www.youtube.com/playlist?list=PLZHQObOWTQDPD3MizzM2x...
https://www.youtube.com/playlist?list=PLZHQObOWTQDMsr9K-rj53...
I would like to be able to read the math in DL papers. (sorry I'm asking for something that it's too broad)
1) How much does this document cover the notations in those papers. 2) When I read a paper and if I am not sure what the math means, does that mean that I did not grok the subject yet, or the math presented in that paper goes beyond the math given in this Matrix Calculus document (assuming I studied well this document).
Note: I come from math and Econ, so the split between practitioner and theorists might be different for CS/ML.
If I could make one request it would be a bit of margin/padding on the left of the body text. Would make it more readable on mobile.