Linear Algebra: What matrices are
nolaymanleftbehind.wordpress.com
nolaymanleftbehind.wordpress.com
If you came up with an isomorphism from matrices to a domain where it made sense to define a different multiplication operation - like maybe placewise multiplication, where
[a b] * [e f] = [ae bf]
[c d] [g h] [cg dh]
then that would be just as valid, but it wouldn't change what matrices 'are'. In fact, because you know how to map matrices to linear functions, it would let you describe an operation to combine two linear functions in a new way and that might lead to some new insight about linear algebra!It's like how, in school you were taught that you can't multiply vectors together. Yet, in shader languages, it turns out that it's really useful to be able to multiply two vectors just by multiplying each component ([a,b][c,d]=[ab,cd]), so they define that as a valid operation.
Are the real numbers "actually" Dedekind cuts or equivalence classes of Cauchy sequences? If we prove that both constructions result in isomorphic objects, what difference does it make? Once equivalence has been established, we're free to adopt either perspective as the situation warrants.
Matrices represent linear transformations whether we want them to or not. As someone else pointed out, the operation you've defined is the Hadamard product[1], which is totally valid but doesn't correspond to the composition of linear transformations.
There are other "products", too, like Kronecker product[2] and the Frobenius product[3], each with the own properties, motivations, and relationships to other parts of mathematics. These are neither good nor bad nor anything else — they just are.
I think it was a misstep for the article to be titled What Matrices Are, because the real idea is that when we think of matrices as representing linear functions then the formula for "standard" matrix multiplication corresponds to the composition of linear functions. It's not just some crazy scheme we invented to torture Algebra II students in high school, but a different perspective on the composition operation that has its own advantages and disadvantages relative to other perspectives.
[1] https://en.wikipedia.org/wiki/Hadamard_product_(matrices)
[2] https://en.wikipedia.org/wiki/Kronecker_product
[3] http://planetmath.org/frobeniusproductAnd yes, understanding that is very important to motivate high school students. Affine transformations provide a good context for that motivation, as well as a good framework for intuiting noncommutativity of multiplication.
From the first paragraph, where he explains the purpose of the article:
> The two fundamental facts about matrices is that every matrix represents some linear function, and every linear function is represented by a matrix. Therefore, there is in fact a one-to-one correspondence between matrices and linear functions. We’ll show that multiplying matrices corresponds to composing the functions that they represent.
And later:
> The connection is that matrices are representations of linear transformations, and you can figure out how to write the matrix down by seeing how it acts on a basis.
If there's a meaningful difference between "a matrix is a representation of a linear transformation" and "a matrix can be viewed as a representation of a linear transformation" it seems largely philosophical and, in any case, tangential to the author's stated goal of explaining why matrix "multiplication" is defined the way it is.
Whether or not matrix multiplication is a "particularly useful operation on a rectangular array in general" boils down to a debate about what is or isn't useful to do with a rectangular array. I'll leave that to other folks with stronger opinions on the matter.
But, a "matrix" comes with the matrix multiplication we all know and love, right? The word doesn't just mean a 2d array. I think of it as a math term, not a CS data structure. More of a class, if you will, data and operations bound together. Not literally and not always, my analogy is imperfect, but, if you asked someone to perform an inner product on a block of numbers, you'd just confuse people if you said "matrix multiply".
For your qualm where people told you "can't" multiply vectors together, what they should have said is "we won't define a multiplication between vectors like we can between scalars" where "we won't" means "we won't for this course." Something I agree that isn't stressed in Math enough in the early years is that it's creation: mathematicians define what they can in order to prove other things, and as long as one can define something that is consistent with other definitions and is logically coherent, it goes. Just like you can define element-wise multiplication, of course it's valid.
Final nitpick: I'd argue an isomorphism between objects A and B is enough to say that A is B, up to some non-isomorphic details, like how they are typeset in an article.
Similarly, there is a difference between an N-element set of numbers and an Nth dimensional vector.
The cross product only exists in three dimensions. And it's not associative (A×B×C gives a different answer depending which order you do it in), which is another thing multiplication usually satisfies.
There are two other not-quite-multiplication operators that I recall seeing. There's an analog of the cross product in two dimensions: (a,b,0)×(c,d,0) = (0,0,ad-bc), so it can be useful to have an operator (a,b)×(c,d) = ad-bc, again turning two vectors into a scalar.
And if the dot product is defined in terms of matrix multiplication by A·B = AᵀB, then you can also define an operator ABᵀ, turning two vectors into a matrix. These vectors don't even need to have the same length.
While true, the Wedge product is a useful concept that generalizes a cross product to arbitrary dimensions and is used in multivariable calculus for proving various integral theorems in high dimension. Here, generalized Stokes theorems apply despite the cross product not being defined. Admittedly, it isn't really a map on the vector space, but the fact that Stokes theorems still hold makes it pretty darn useful to me.
So just as every matrix is a representation of a linear transformation of vectors into other vectors, with matrix-vector multiplication corresponding to function application, it is also true that each vector in a vector space represents a linear transformation that turns vectors in the space into a scalar, with vector-vector multiplication in the form of inner products corresponding to function application. The converse is also true: every linear functional on a vector space can be represented by a vector in the space.
This last insight is known (in various forms) as the Riesz representation theorem and holds not only on finite inner-product spaces (i.e., vector spaces on which an inner product is defined) but also on Hilbert spaces (complete inner product spaces, whether finite or infinite). It turns out to be quite powerful.
This is what a 3x3 matrix is: https://dl.dropboxusercontent.com/u/364079/WhatAMatrixIs.png
Just one of the many reasons you don't want to use R^4 to describe a matrix is because, for an orthonormal basis, A^-1 == A^T. The inverse and transpose are the same thing. That doesn't work in any arrangement except a rectangle.
(edited for spelling)
The documentation is extensive, but unfortunately too much of it takes the form of the same garbage code.
I was taught discrete math from the building blocks of matrices and it seemed as useful at the time as learning driver's ed in a license plate factory.
However, I don't think in the slightest that throwing CS buzzwords such as "REPL" and "hackable" at the problem is the way to go.
Fill inn these boxes and get the solution is not hackable (unless the programmer seriously failed at validating input.)
The filling in of the boxes is most fairly compared to finding the typo in the complicated regex, or the file with incorrect permissions on the web server
This style of learning is what dominates US schools now from basic arithmetic all of the way up through advanced calculus.
The point made in the parent comment is that teaching math often fails in its intent, to provide students with the rewards of mathematical insight and ability to reason.
Instead basic math education results in a rift between those who "aren't good with numbers" and, well, masochists.
What is it about math education that so abhors references to applications?
People learn about mammals in biology and don't ask for references to applications. Why is it different in math?
This thread began with a parent article on matrix algebra, which has many, many, exciting real-world applications.
That said, there is an audience for whom the beauty of the subject is enough. Now when I hear "Galois theory" or "Navier-Stokes" I just think "I don't know what that is, but I want it!"
In your weekly assignments (but not exams) the teacher may even fail a problem if you only write the correct answer, but fail to show how you got there.
Is that the kind of thing you are asking for?
For linear transformations, one possible avenue is start with the notion of derivative. If we take a real-valued function, its derivative at a particular point is just a single number, which represents the function's rate of change. But what if we have a function that accepts, say, two real numbers and outputs three? It turns out that the natural generalization of "derivative" to such functions is a rectangular array of numbers:
dy1/dx1 dy1/dx2
dy2/dx1 dy2/dx2
dy3/dx1 dy3/dx2
If we know these numbers (and nothing else), we can linearly approximate the values of a function near a particular point, with at most quadratic error.Now let's say we have two functions. The first one takes two numbers and outputs three, and the second one takes three numbers and outputs four. If we compose them together, can we find the derivative of the composite from the two simpler derivatives by some kind of chain rule, like the one we have for ordinary real-valued functions? It turns out that yes, we can, if we replace the product of two numbers with the product of two matrices (defined in a particular way).
Now it's easy to explain what linear transformations are. They are just multidimensional functions whose derivative (matrix) is the same at every point. They are just like one-dimensional linear functions, whose derivative (number) is the same at every point. (For convenience, people also say that every linear transformation must take the point (0,0,...) to the point (0,0,...), so that matrices correspond one-to-one to linear transformations and vice versa.)
If you want to work with linear transformations effortlessly, there's a lot more intuition to develop, but this should serve for the basics.
Ignoring the fact that you need the function's values actually at the particular point in question, you do in fact need to know something else: you need to know that the function in question has sufficient regularity (enough smoothness) to allow an application of Taylor's theorem.
http://www.corestandards.org/Math/Content/HSN/VM/
CCSS.MATH.CONTENT.HSN.VM.C.6 (+) Use matrices to represent and manipulate data, e.g., to represent payoffs or incidence relationships in a network.
CCSS.MATH.CONTENT.HSN.VM.C.12 (+) Work with 2 × 2 matrices as a transformations of the plane, and interpret the absolute value of the determinant in terms of area.
Math education gets a bad rap in the USA, because anything that students don't remember learning, is claimed to have never have been taught. And that is just not true. Math teachers aren't as dumb as people pretend they are.
I have seen the standards for high school math teachers and trust me, they are low. Most of my classmates in undergrad (which was supposed to be an excellent program for math teachers) are now high school teachers. Their complaints about how much they hate basic linear algebra and "just want to be done with it so they can go teach high school" are still ringing in my ears.
The article is right in my case, I didn't learn the intuitive understanding at first, and it could have been taught that way. It is also well written, for people that know math, but I also feel like the article describes math using more math, and that the point could be better made with a picture or two. There's something about just seeing the correspondence between matrix rows, or columns, and the axes of a space, or transform, that finally helped it all sink in for me.
Everyone just tells you, "oh now this can be a matrix." Nobody tells you why matrices are useful, or how we ended up with them. Just, you know matrices use them.
As anecdotal evidence, most math teachers I had in college were at best uninterested and at worst disdainful of applying the math to the real world beyond theorems in a blackboard. Most physics teachers, however, in introducing us to new concepts made an effort to show as soon as possible how the definitions we made were inspired by real world problems and in turn simplified or helped create new physics. This made it much easier to appreciate the pure math itself.
The response is often, No we do X because of Z. Which often you're learning Z. So now you just feel lost and confused. Its not 2 classes later until you'll learn A maps Y to Z and X is actually an operation of A so it applies to both.
I guess it's just me but the joy of math has always been its tangled relationships to itself.
You mentioned opengl and it's a perfect example of why you need to learn both. The Model matrix is column-oriented - if Z is up in the model but Y is up in the world, you need a Y unit vector in the third column of the matrix. The View matrix is row-oriented - each row is an axis of the camera's orientation. The columns perspective also makes it easier to understand homogeneous coordinates.
Mathematics never pays attention to what objects ARE, but rather what they DO.
Matrix is not defined just by operations like a ring is, but also by structure - you have N independent axes in a specified space.
> If you can do numeric arithmetic on it, it's a number.
>you have N independent axes in a specified space.
Only if you have an operation that maps the matrices to that space.
Similarly, mathematicians study objects by looking how they act on other objects. So if it turns out that two differently defined objects have the same actions on other objects, we call them isomorphic and don't really distinguish between them.
So for instance, we say that the set of rigid motions that preserve the triangle and it's orientation is the same as the set of permutations of the roots of say: x^3-3x+1 even if the two sets are absolutely not defined in the same way.
Hope that makes some sense.
For example, if we are considering sets and functions between them, we generally don't care about the exact names of the elements of the set, only the fact that they are distinct elements. The important aspects of a function in this case are properties such as injectivity, surjectivity, etc, not that it sends one particular element of one set to another particular element of another set.
Another example is in linear algebra: we really care about linear transformations on an abstract vector space more than we care about what that linear transformation looks like relative to a specific set of coordinates.
This point of view is espoused in category theory, where the important information is carried in the morphisms between objects, not really the objects themselves.
http://betterexplained.com/articles/linear-algebra-guide/
But to me, it's better to start with a context, a purpose for using matrices & linear algebra first, and learn what and how to use matrices in that context. The contexts that helped me included 3D graphics/games and later circuit simulations.
http://www-groups.dcs.st-andrews.ac.uk/~history/HistTopics/M...
This is in some sense the process all math students go through. The formulas for computing determinant and multiplying matrices look really complicated and it feels like a mystery as to why it works at all but then linear algebra explains all of that slowly.
Not trying to be snarky, just an observation.
https://www.maa.org/external_archive/devlin/LockhartsLament....
(1) They are "nice" operators which carry properties most functions can only dream of having: f(v + w) = f(v) + f(w), and f(av) = af(v). From these properties you can develop a very rich theory (along with the properties of vectors). These are the sorts of properties that we want all of our functions to have when we are young and first learning algebra - how many prealgebra students wish they could simplify (a + b)^2 to a^2 + b^2?
(2) A linear approximation is the first useful approximation for most behaviours, and at a small enough scale almost anything looks linear. Once you've made this approximation you get to exploit the properties of (1)