Linear algebra for programmers
coffeemug.github.io
coffeemug.github.io
For more info about SymPy, see section "VI. Linear algebra" in the SymPy tutorial I wrote https://minireference.com/static/tutorials/sympy_tutorial.pd... (also available as notebook https://github.com/minireference/sympytut_notebooks/blob/mas... )
Aggressively taking the least shortcuts possible is the fastest shortcut.
To me, the didactic style is what matters, not the intellectual prowess of the mathematician in question. For this topic, and the depth to which I wanted to expose myself, I've found Pavel more accessible than anyone else.
(Watch them at 3x, and then at the end of each one, reward yourself by setting it back to 1x, and Gil Strang sounds like he's on a Robotussin bender.)
I can sorta kinda get the theory, but every demonstration involves moving an arrow around which is.... not something I need to do frequently. So I'm not sure how I actually apply linear algebra to solve actual problems.
I'm a software developer and I know it's useful I just don't get where to use it - and I'm struggling to actually understand the different operations, purpose of the dot product, etc. I have a decent base for basic stats and calc, both of which I can "conceptually apply" near daily for understanding how things work.
3Blue1Brown is helpful, but I just kinda go "yeah I guess that looks right" without knowing what to do with it.
EDIT: Thank you!
You can also go in the direction of cryptography; here's, for instance, a really excellent LLL tutorial that builds on Graham-Schmidt: https://kel.bz/post/lll/
The specialization as a whole starts simple, and works its way up to PCA.
https://www.coursera.org/specializations/mathematics-for-mac...
As for why it comes up so much, it concerns itself with solving systems that look like y = Ax + b, where the Ax term works similarly to multiplication in 1-D. The point is these simple equations are ones we can actually understand! Everything else is too hard.
But there's a trick we have for everything else: if you have some y = f(x) where f is super complicated, you can differentiate. The derivative of f at a point x_0 is the best linear approximation to f. i.e. f'(x_0) is the best matrix A such that y ~= Ax + b near x_0. Now your problem is linear and you can understand it (locally)! Then you can integrate your local solutions into a global one.
The purpose of dot products is that they let you talk about things like angles, lengths, and projections. The point is you learn how it works for arrows and shadows and stuff, and figure out some equations that hopefully make intuitive sense in 2- and 3-D, and then it turns out those equations work in higher (even infinite) dimensions too.
Projections are useful because they let you break vectors down and build them back up, and hopefully the broken down version is easier to understand. Understanding projections in high or infinite dimensions gives some intuition for things like the Fourier transform, where you project a function onto simpler waves, maybe study how a system reacts to those waves, and then use that description to build back how up the system reacts to your original function.
Angles give one way to measure closeness. If you have some machine learning model that figures out a way to map text into a 50,000 dimensional space, you might be able to do it in a way where two sentences are intuitively similar if they are mostly pointing off in the same directions, so if the angle between them is small.
So tl;dr the idea is you learn some geometry with arrows and all that, you figure out some equations from that geometry, and then you realize that those equations and that geometric intuition work anytime you have a linear (i.e. f(ax+b) = af(x) + f(b)) system. Calculus gives you ways to turn non-linear problems into linear ones, so you will find examples of linear systems everywhere.
"Coding The Matrix: Linear Algebra Through Computer Science Applications" https://codingthematrix.com/
It used to be on Coursera too, has video lectures and a book to go with it.
It has companion code in Julia and Python, and addresses many important applications, such as convex optimization where the authors' previous book is really famous.
It turns out dot products, matrix multiplication, etc. are fairly common in projecting a 3D world onto a 2D display (flight simulators, for example).
Edit: I should mention that this is a good way to learn to appreciate what modern shaders are doing. You could jump into those as well, but I think it's a little harder to get your first triangle on the screen. All this assumes you're comfortable writing C, of course.
But if you really want practical applications, you probably want to read a book on numerical analysis, which is where linear algebra really finds its glory.
A simple way to think about Linear Algebra is as a set of rules i.e. functions (aka Matrices) applied to input points (aka vectors) to give output points in a vector (i.e. multidimensional) space. Thus multi-valued inputs give multi-valued outputs. The key is to keep both the algebraic manipulations and what they may mean in the geometric interpretation simultaneously in mind. The basic example of solving a system of linear equations in two variables will cement this idea when you map it to the XY coordinate system.
You might find the book Practical Linear Algebra : A Geometry Toolbox by Farin and Hansford useful for further study.
Machine Learning needs some things (mainly "easy" things like matrix multiplication and convolution with the occasional truncated SVD), Modeling of Physical Systems needs other things (PDE solvers), Computer graphics different stuff (mainly focusing on small rotation/transformation matrices applied to many many points), Nonlinear Optimization yet another subset (solution of large sparse systems), and that's before you get to signal processing and statistics.
The main commonality is once you have the basic terminology down and understand that if you're explicitly inverting a matrix you're probably doing it wrong, you should just use the existing highly tuned libraries for your use area (at least until you decide it would be cool to try and beat them). Once you need to go beyond that you're more into the realms of matrix analysis and structure exploitation, and firmly have at least one foot in the math camp.
And, truly understandng the subject will pay you back in spades. Don't go for fluff. Linear algebra isn't some Facebook post. Truly go through the book line by line and really understand it down to the nuts and bolts. Suffer a little.
However i fear it is lost on the current generation which is always looking for that quick/short article/blog/video/etc. which will teach and make them understand the subject as painlessly as possible. Everything must be "Fun and Enjoyable" (i don't even know what that means anymore!). They have forgotten the maxim "There is no Royal Road to Mathematics"(or any other scientific field). The amount of people (even on HN!) who espouse disdain for Textbooks is astonishing. By definition, the process of learning involves diving into the unknown and hence there will be discomfort/difficulties/effort/time needed and thus is not going to be easy. But the student doesn't want to put in any effort at all; everything is the fault of the Teacher(can't teach properly)/Teaching method(i am a visual learner)/Too rigorous/Mathematical/I have ADHD/ADD/Autism/Whatever.
The real tragedy is that used Books are now so easily/cheaply/universally available that there is no reason not to have one's own library of books on subjects of interest; it is "food" for the Mind.
I don't see how 3 can be a function from this example. "3*" (partially applied multiplication by 3) looks more like it.
Matrices and vectors as functions? Yeah, if the argument is within bounds. That makes it just an indexing operation.
(I guess one can view 3 as a one element vector but that sounds like a degenerate case)
Or maybe I'm missing something...?
But I think that the technical, mathematical way to think about it is:
The monoid of linear functions L:R->R is isomorphic to the monoid (R, *)
Meaning, the structure of 1x1 matrices under multiplication is exactly the same as the structure of real numbers under multiplication.
Seems trivial but among other things it implies associativity, which is not quite trivial for larger matrices.
Intuitively, natural numbers come from and are defined by counting, and that implies that "3" means inherently that something (could be anything) happened or was repeated three times. For example, if you have three apples, that means that you can identify one particular apple that you have, then do that again, then do that again.
Adding a unit to a number is like adding further information on what it is that is being repeated. Three pairs of apples? You just invented the number six!
The meaning of doing something three times (most abstractly: applying the successor function) is already inherent in the meaning of three, so multiplication isn't something that has to be added on top. It's already in there.
It really helps build intuition beyond the typical teaching methods.
It will really help connect the dots.
https://youtube.com/playlist?list=PL0-GT3co4r2y2YErbmuJw2L5t...
also, why say these 2 things?
>If you forget how matrix-vector multiplication works, just remember that its definition flows out of the notation.
>Another way to think of matrix-vector multiplication is by treating each row of a matrix as its own vector, and computing the dot products of these row vectors with the vector we’re multiplying by. How on earth does that work?! What does vector similarity have to do with linear equations, or with matrix-vector multiplication?
But you just told me to plug into the definition? Which is a dot product.
Pretty incoherent.
it hasn't gone through a spell-checker, with words in there like hight and strage
Moreover, here as in most other writings about linear algebra that I have seen, there is the very bad habit of describing the more complex operations as being composed from dot products.
On modern CPUs, dot products must be avoided, because like all reduction operations they consist of one chain of dependent operations, so their speed is limited by the latency of the fused multiply-add operations, instead of being limited by the much higher throughput of the FMA operations.
When vectors are multiplied with vectors, there is no alternative for dot products, so the only way to accelerate them is to reorder the operations into a tree, to be able to overlap a part of them, which are independent.
When arrays with more dimensions are multiplied, e.g. matrices with vectors or matrices with matrices, the multiplications correspond with nested for loops, i.e. 2 nested loops for matrix-vector multiplication and 3 nested loops for matrix-matrix multiplication.
The nested loops can be reordered arbitrarily. In each case, one of the possible loop orders has in the innermost loop the computation of a dot product.
This is the only order mentioned in the parent article and in most other linear algebra manuals. However, this order is exactly the worst possible computationally.
For 2 or more nested loops, there is always another order where the innermost operation is a so-called AXPY operation (the BLAS function name). AXPY means the scalar A multiplied by the vector X Plus the vector Y, with the result stored in Y. AXPY operations are always better than dot products, because the FMA operations are independent, so they can be pipelined, but for 3 or more nested loops there are better orders, where the innermost loop needs much less load and store operations than for AXPY.
For 3 or more nested loops, there is always an order where the innermost operation is a tensor product of 2 vectors. This is a more attractive operation than both AXPY and dot products. If the tensor product of 2 vectors with N elements is stored in registers, then its computation needs N+N loads from memory, but N*N FMA operations, so if there are enough registers so that N>2, there will be much more FMA than loads, allowing the full utilization of the execution units of a modern CPU.
In conclusion, matrix-vector products must not be described as being composed of N dot products that are executed separately for each element of the result vector, but as being the sum of N AXPY operations, which are accumulated into the result vector.
Similarly, a matrix-matrix product must not be described as being composed of N^2 dot products, one for each element of the result matrix, but as being the sum of N vector-vector tensor products, which are accumulated into the result matrix.
Such descriptions would be much more useful in practice, where the definitions based on dot products are just a hindrance.
I wouldn't know for sure because it all whooshed about 60 feet over my head
No offense but I stopped reading there. Too many software developers have this weird superiority complex when it comes to math. When they struggle with math, I've seen devs criticize everything from naming conventions to curricula to it being "useless". A lot of them seem unwilling to acknowledge that math is sometimes... simply hard.
If your attitude to math is "I don't need any of this useless stuff, so I'll zone out", then I kindly suggest you first at least try to learn Linear algebra for mathematicians first.
Software is a big field and there's room for folks with different interests and talents.
I followed a book in the 90's to create a flight simulator from scratch. Besides learning Bresenham's line algorithm, I learned a lot of linear algebra.
Probably this book: https://archive.org/details/build-your-own-flight-sim-in-c-d...
I always recommend “Ray Tracing in One Weekend” [1] series of books for some light coverage of linear algebra/geometry/graphics.
I was 13 at the time, so I struggled with the math (this was pre-internet so I couldn’t just search stuff). I’d never seen matrices and wasn’t able to figure them out from that book. I distinctly remember hitting a wall in the chapter on lighting because it mentioned taking the dot product of two vectors, but I didn’t understand why that would give you the cosine of the angle between them because I’d never heard of the dot product before.
Edit: right at the bottom of page 405. I reread that sentence so many times.
The point is to teach it to people who never understood its value but now want to learn, and to do that it helps to relate to their experience.
However, the naming conventions and syntax overloads/ambiguities really can be a pain as a non-mathematician.
I get it. Everything has warts, and the mathematical syntax is as much for thinking and experimentation as it is for communication, and verbosity hinders understanding. I'm also aware computing has its horrible dark corners. But from the outside, it seems like maths practitioners often oppose any attempts to make it any clearer, even to people in other mathematical subfields.
Papers leave variables undefined because everyone working in that subfield is expected to just know what that variable means in that context, possibly derived from on some book that everyone in that field knows about so nobody feels the need to specify it. I mean just a stupid but simple example, I learnt maths from an applied engineering school so the imaginary unit was j. Always j. Never specified that it could be anything else nor was it ever specified anywhere that j was the imaginary unit. Then I start exploring the topic outside textbooks and everyone's using i. Again, no indication that this is the imaginary unit, it's just a given. Obviously in this particular case it's easy to notice that it's been swapped over, but there are plenty of other much more subtle ambiguities that can make trying to understand anything a real slog.
That said, I think and hope that I may in the future feel differently as my skills improve. My theory is that if everyone always added in that foundational knowledge to each paper etc it would make everything really verbose and make trying to get to the point of what you're saying a real slog for the author and the experienced practitioners. Being able to be consise means you get to the heart of the new stuff quickly without having to slog through a bunch of "C is the set of complex numbers a+bi where a and b are in R and i^2 = -1" first.
Except once someone did that, you could literally just cite their paper or book.
Does the "es" in ESLint mean it only works for ECMAScript? Or does the "ES" mean something entirely different? The homepage doesn't say. Are eslint, ESLint, and Eslint the same thing? Capitalization usually matters in software after all, but nobody is consistent here, not even eslint.org. Why do are ES6 and ES2015 used interchangeably? That's unnecessarily confusing.
All of this is far more confusing than "i" vs "j" for sqrt(-1).
> criticize everything from naming conventions
A programmer criticizing naming conventions. a.k.a another Tuesday.
But here is the thing. As a software developer, there IS a "higher being" aka a designer for the programming language or library and they try and try and try their best to maintain "rhyme and reason". With math? There is no such thing, or rather, everyone pisses in the pool of math notation but nobody wants to admit that and so you get confused people, people who mistakenly look for the "rhyme and reason" and they find none, they find no pattern.
So the first lesson about math notation that you need to learn is that there is no such thing. There is this chicken scratch that makes it easier to write on the blackboard or on paper and that is about it.
Math isn't "hard..." it is abstract.
My hint for something to play with is that basic linear algebra applies very directly to graphics, rotation matrices and so on. If you know how to multiply a matrix with a vector, you basically know what you need to render basic line-art 3D graphics. May want to look into dot and cross products as well as vector projection, but it's fairly basic all of this.
But why?
There doesn't seems to be a lot to learn about applied linear algebra (in the sense discussed here) by implementing a rasterizer.
But there is plenty of LA above and below that (in the scene management, in the shaders).
Is there a githubrepo/tutorial for how linear algebra is used for a very small model just to demonstrate how that allows it to "learn"?
I've got the calc, I just don't understand what the matrix multiplication "does"
https://www.coursera.org/learn/machine-learning
I cannot speak to the author of the content of this github repo, but it appears they have completed the course and included all of the solutions here. It might let you jump right to what you're looking for.
https://github.com/greyhatguy007/Machine-Learning-Specializa...
[1] https://www.youtube.com/watch?v=VMj-3S1tku0 [2] https://github.com/karpathy/micrograd