Inside the Matrix: Visualizing Matrix Multiplication, Attention and Beyond
pytorch.org
pytorch.org
* "The Essence of Linear Algebra," by 3Blue1Brown: https://www.3blue1brown.com/topics/linear-algebra
* The popular introductory course taught by Gilbert Strang: https://ocw.mit.edu/courses/18-06-linear-algebra-spring-2010...
I cannot overstate how good of a teacher he is. Things click into place while watching him teach.
Reminds me of that scene in the movie, CONTACT, where the aging mathematician solves the alien primer by realizing the documents assemble into a 3D cube and the decryption happens when the squiggles on opposite sides of the cube are overlaid resulting in clear text.
There are all kinds of fascinating places where you can gain mental leverage by thinking in higher dimensions. For example, the definition of a monoidal category, which includes various equivalences (or for a strict monoidal category, equalities), can be seen as telling you about the existence of certain 3-dimensional "sheets", 2-dimensional slices of which are equivalent (or equal) ordinary functorial string diagrams[0]. This is just a higher dimensional extension of the fact that chaining 1-dimensional slices of functorial string diagrams give you particular paths in an ordinary commutative diagram. see Marsden [1] for more on that.
Unfortunately the computer tools for generating and manipulating these kinds of topological constructs are in their infancy, which is probably why they aren't used much by mathematicians.
[0]: https://twitter.com/nathanielvirgo/status/126201964172083200...
Sure it looks cool, but mostly because it looks magic, which is the opposite of what's you're supposed to do when illustrating mathematical concepts…
The typical textbook illustration using 2x3/3x2 matrices is much, much clearer than this 32x24 … 64x96 mess. The 3D idea is interesting, but why spawn such an insane amount of elements in your matrices?!
Colorized Math Equations[2] has the same problem where people see it and go "Colors! English language! This must be so much more easy to grasp than math! I feel enlightened for having seen this!" But feeling enlightened is very different from being enlightened and it just doesn't hold up. I've found people retain very little understanding if they aren't already familiar with the concept.
[1]: https://byorgey.wordpress.com/2009/01/12/abstraction-intuiti...
[2]: https://betterexplained.com/articles/colorized-math-equation...
EDIT: The "three-dimensional operation" perspective no doubt comes from writing matrices as rectangles, but this is far from the only representation of them. If the vector v = [a, b, c] is shorthand for v = a x_hat + b y_hat + c z_hat (explicitly a sum of basis vectors), then we can write a matrix with a similar set of basis vectors: m = [[a, b, c], [d, e, f], ...] = a x_hat x_hat + b x_hat y_hat + c x_hat z_hat + ... . There's nothing "rectangular" about this any more than a polynomial (as a sum of monomials) is "rectangular". The details then shake out of how (x_hat y_hat) multiplies with (y_hat z_hat). The rectangle is just a mnemonic.
DOUBLE EDIT: In the above sense, multiplying two matrices is more like a convolution -- the x_hat x_hat term of the first matrix multiplies every term of the second, we just know most of those terms will be zero (the product with any term that doesn't start with an x_hat (e.g. y_hat z_hat).
The animations really helped me to understand what eigenvectors, eigenvalues, linear transformations, determinants etc are
It would have been more intuitive to show every element in output matrix corresponds to a dot-product of row/column vectors from input matrices, the animation doesn't even highlight those corresponding vectors clearly..
If it's useful to someone working in a very complex environment where these visualizations are necessary to help tease out some subtle understanding, then that's great.
But really, this part is all you need to know about the article:
> This is the _intuitive_ meaning of matrix multiplication:
> - project two orthogonal matrices into the interior of a cube
> - multiply the pair of values at each intersection, forming a grid of products
> - sum along the third orthogonal dimension to produce a result matrix.
This 1. Isn't intuitive, and 2. Isn't the "meaning".
This gets to the why perfectly. We all understand how to navigate a 3D space intuitively. If the math doesn't tie into that, it may as well be wizard nonsense.
1. For a linear function f, its matrix A for some basis {b_i} is the list of outputs f(b_i). i.e. each column is the image of a basis vector. For an arbitrary vector x, the matrix-vector product Ax = f(x).
2. For two linear functions f,g with appropriate domains/codomains and matrices A,B, the result of "multiplication" BA is the matrix for the composed (also linear) function x -> g(f(x)). For an arbitrary vector x, the product (BA)x = B(Ax) = g(f(x)).
This tells you what a matrix even is and why you multiply rows and columns in the way you do (as opposed to e.g. pointwise). This also tells you why the dimensions are what they are: the codomain has some dimension (the height of the columns) and the domain has some dimension (how many columns are there). For multiplication, you need the codomain of f to match the domain of g for composition to make sense, so obviously dimensions must line up.
Respectfully disagree. A matrix has basically nothing to do with "living on the surface of a cuboid". It's like saying FOIL is the "why" of binomial multiplication -- the "why" is the distributive and associative properties of the things involved, FOIL is just a useful mnemonic that falls out.
There are already rich geometric interpretations which provide useful intuition for that generalizes, rather than just demonstrating mechanical details.
In short you need to understand: vector, linear combination, cross product, partial derivative, chain rule and finding global minimum. If you have basics of linear algebra it’s easy to grok this video.
This is the equation of a neuron is you squint.
So if you chain a bunch of neurons, you are basically drawing a bunch of lines to test whether points belong or not.
With enough lines you can approximate any shape, like a circle.
What neural network do is given enough examples, it finds the lines that are needed to separate the points to give the appropriate label.
It builds a simple CNN and ends with a simple example of how multiple ReLU activation functions can approximate arbitrary curves.
A = UH
Each of U and H consists of real and/or complex numbers and all real numbers if A consists of all real numbers.
Here U is unitary which means that for any n x 1 vector (real and/or complex) x, Ux is the same as x except is rotated and/or reflected and the lengths
|x| = |Ux|
that is, U does not change lengths or distances and, thus, is a rigid motion (rotation, reflection).
For H, for x in a sphere S, the set of all Hx is just an ellipsoid.
So, A = UH where U is a rigid motion, rotation, reflection, and H converts a sphere to an ellipsoid. The H is said to be Hermitian (there was a mathematician Hermite). Right, an ellipsoid has mutually perpendicular axes, and those are the eigenvectors of H.
My favorite result in linear algebra, and with a fairly short and simple proof.
I think the closely related SVD is more commonly used though, and Wikipedia has a nice animation https://en.wikipedia.org/wiki/Singular_value_decomposition