Matrix calculus for deep learning part 2
kirankamath.netlify.app
kirankamath.netlify.app
I _much_ prefer to reduce a given matrix expression into einstein summation convention, at which point all of the "regular" calculus rules just work. You can bash it out from this point on.
For example, consider the case of `x^T x`. We are told from matrix calculus that this is `2x`. To do this using summation convention, we first write it in terms of coordinates. We will have:
y = xi xi [summation over i implicit]
dy/dxj
= d(xi^2)/dxj
= d(xi^2)/dxi * dxi/dxj [chain rule]
= 2xi delta(ij) [all xi independent, dxi/dxj = dirac]
= 2xj [summing over i]
dy/dx = 2xAlso, for this trick to work on matrices, you need two indices.
It's when you take derivatives of vectors & matrices by other vectors & matrices that things get "interesting".
[1] http://www2.imm.dtu.dk/pubdb/views/edoc_download.php/3274/pd...