Matrix Calculus: Calculate derivatives of matrices
matrixcalculus.org
matrixcalculus.org
[1]: https://papers.nips.cc/paper/2018/file/0a1bf96b7165e962e90cb...
[2]: https://www.youtube.com/watch?v=IbTRRlPZwgc
[3]: https://compcalc.github.io/public/laue/tensor_derivatives.pd...
A good book on differential geometry will also probably start with an overview of tensor calculus.
[1] https://ojs.aaai.org/index.php/AAAI/article/download/5881/57...
You want to solve Ax = b approximately. So, minimise the two-norm |Ax-b|, or equivalently, |Ax-b|^2, or equivalently (Ax-b)ᵀ(Ax-b) = xᵀAᵀAx - 2xᵀAᵀb + bᵀb.
How to minimise it? Easy, take the derivative wrt the vector x and set to zero (the zero vector):
2AᵀAx - 2Aᵀb = 0, so x = (AᵀA)⁻¹ Aᵀb.
(Note: that's the mathematical formulation of the solution, not how you'd actually compute it.)
For example, if you recall that (due to stokes theorem) the divergence is minus the transpose of the gradient, you can solve many variational problems that way. The canonical example is: find a function u whose gradient is a given vector field F. This is an over-determined system without a solution. Applying the above trick, you compute the divergence of both sides to obtain Poisson equation Δu=div(F), that you can solve. This is equivalent to finding the least-square minimizer of the energy E(u)=∫|∇u-F|²
Even plain old multiplication and division, and even addition and subtraction have stability and efficiency problems on floats, which don't appear in symbolic solvers.
Also the output is pretty gross, wish it had an option for a statically typed language.
Also wtf? It doesn't have sqrt? Have to write it in power form..sigh
Also seems to also not grok values with decimal points as the power, so you have to write it as a fraction..sigh..why math people..why?
Also why can you only select the output for 1 value at a time? For instance if we have a 3d position, we want gradient for x/y/z not just 1 of them..
In an example they use X' for transpose and it works.
But if I try A' * A, it either becomes just A * A, or it shows nothing and says " This 4th order tensor cannot be displayed as a matrix. See the documentation section for more details."
The documentation also doesn't show how to enter the transpose operator.
Doesn't this cover all examples presented?
It should be derivatives "with" matrices, not "of", in my mind.
Not that it matters in practice, but ... if there's one field where precision of language matters, it should have been mathematics. So it bothers me.
But if that collection is plugged into an expression, so as to denote a collection of expressions when result when the individual element vectors are plugged into the expression, and the resulting expressions are differentiable, and you then differentiate each of those expressions and then collect the result back in the same order and represent it as a matrix, then that's fine.
This is what happens when we talk about "matrix differentiation", we mean "differentiation with matrices". But the matrix itself is not a derivable object. Only the expressions upon which it confers its notational semantics have a derivative.
I think you're arguing that a matrix with constant real entries isn't differentiable with the standard derivative. This is true of course, but hasn't anything to do with vectors.
A derivative is a derivative of a function (or a function like object, say a functional). The phrase "the derivative of a matrix" keeps it unclear whether "the matrix" is the function, or whether "the matrix" is a variable, an element in the domain of the function.
Bereft of that information, the phrase is as meaningful as "the derivative of 4.2"
I have a similar annoyance with the notation for probability. I think stuff like `p(a)` is a historical mistake. We would have been better served with a notation that would make clear the distinction between a domain as an entity in itself, the set of possible values in the domain, and the action/event of selecting one or more values from that domain.
But oh well. The naming problem is one of the hardest problems in science after all :)
This is 0 for any matrix of constants, as you can see in the example.
The derivative of x² is 2x (where x is a scalar)
The derivative of vᵀv is 2v (where v is a vector)
There's no differential equation.
http://www.gatsby.ucl.ac.uk/teaching/courses/sntn/sntn-2017/...
I used matrix derivatives in grad school a lot (my dissertation was on nonconvex optimization) and I'm not sure I've ever needed an entire textbook. Matrix calculus is just plain calculus with a few extra rules to generalize it to matrices/vectors.
(to be effective, you do need to know the rules of matrix algebra however)
On the other hand, if you are already familiar with calculus and linear algebra, then most of materials are just straightforward derivations using theorems you already know, so you might not need a textbook. I liked this textbook because it took mysteries out of formulas I was taught but never told about derivations, and because it was actually a quick read.
There are some limitations though. The book stops at Hessian matrices and matrix derivatives of vectors or something like that, because chain rules start breaking down and one needs multilinear algebra for higher order derivatives. However, for deep learning you don't need to worry about it. The book will teach you enough to derive the gradient of convolution as realized in matrix products.