Backpropagation algorithm visual explanation
google-developers.appspot.com
google-developers.appspot.com
For those without a math background, the notation is very opaque. A far better explanation is to explain it numerically with simple examples.
For example, have two bits of training data:
input -> output
1 -> 0
0 -> 1
And a simple network with zero hidden nodes and train it. By hand...
Then add another bit of training data:
0.5 -> 1.5
Notice that it is now impossible to fit the training data exactly, however many training iterations we do. Now add a hidden layer with one or two nodes. Now we can perfectly fit the data, but show that depending on initialization weights we might never get there through gradient descent. Nows the time to mention different types of optimizers, momentum, etc.
Anyone targeting their work at non-experts should explain even what seems like trivial notation to them since they can't know what other meanings the reader may think the notation holds.
Computer languages benefit from the fact that poorly designed syntax can be deprecated (not in all cases, e.g. C++) by introducing new features to the language.
Notation in math never advances in the same way for some reason.
Specifically for the superscript case, the vast vast vast majority of the cases, it will be obvious whether the superscript notation means exponentiation or indexing. When there are ambiguities, what the person explaining the equation should do is clarify the ambiguity.
It doesn't do the world any good to make the entire engineering world learn (relatively) obscure programming languages in order to be slightly more clear (and let's not forget, much more verbose) when writing down simple equations, when all one has to do is clarify a couple of ambiguities when writing down equations.
Let me make it very clear and say this: equations are meant to explain things. They are not standalone pieces of code that you can copy and paste into a REPL. They are tools to explain how something works. You should always have accompanying text that explains what all the variables in the equation mean and clarify unclear notations.
I'm relatively sure that they teach basics derivatives, functions etc. in high school in every country.
Here in Finland the students are divided into "long" and "short" math. Both will learn basics of calculus and derivatives at least.
You're wrong, they don't teach derivatives in the HS curriculum in Poland, even when you take advanced math in HS.
But you are going to hear it anyway if you are going to study.
I believe it was covered at A-level (18 year old) but you could only pick three or four subjects for A-Levels at the time, so you had to be selective about the subjects you picked depending on what you "want to be when you grow up", and what you thought you were good at so that you got good enough grades to go to a uni you liked.
I think the intuition is only half-transmitted with just the notation.
I do think working through an example sort of completes the intuition.
For me it depends on whether I'm in a passive or active state of learning.
If I'm sitting down on a Sunday afternoon reading the news, backpropagation is going to make zero sense to me.
But, if I'm actively working on a problem, it's much more useful to realize that this is no different than using gradient descent for linear regression or even minimizing a quadratic.
At that point, it becomes just a mechanical calculation (the fact that the resulting gradient looks more intimidating is irrelevant).
And for me, realizing that it's no different than taking the derivative of a quadratic actually makes it more digestible than these fancy animated tutorials.
If you understand that already, then you are really not the target audience of blog posts like that.
The key idea of backpropagation is that at each layer, you only ever need the derivatives of the loss function w.r.t. the parameters of the layer, and the Jacobian-vector product with the derivatives of the loss function w.r.t. the layer outputs. You never need to compute the jacobians explicitly and you never need to do those high dimensional matrix-matrix multiplications.
These are not complicated ideas but they involve a combination of software design, calculus, and linear algebra that would probably not be obvious to the average CS undergrad.
I'm not sure if it's the same for others, but I don't find bare text descriptions with formulas particularly useful. Mathematical notation on a page is great for rote application of rules and computation, but by itself does not easily communicate an intuitive understanding of the system the math represents. I have to work very hard to build up mental pictures of systems described by just notation, and those mental pictures often have to move in complicated ways as well.
The relationship between maths on the page and the systems they describe is the same as seeing musical notation on a page and hearing a full orchestra. One is a dry accounting of the facts involved. The other is moving and powerful in its richness and immediacy, a living thing that defies easy communication beyond the experience itself.
Demonstrations like this show you the maths _and_ build up a picture for you at the same time. The result of that is that you can communicate a very powerful idea (e.g. backpropagation) very precisely, intuitively and quickly.
Very much worth a five minute scroll for me; YMMV!
In English we read from top to bottom. Data flows (be it equations or flow charts) typically follow the same convention, so we can read articles in in a coherent way. Even trees (both data structures and decision trees), grow from their roots downwards (so, against their original biological metaphor). At least most of researchers write neural networks from left to right, consistent with English.
More of this point: https://www.reddit.com/r/MachineLearning/comments/6j28t9/d_w...
- top->bottom is not compatible with left->right - "back" propagation is "back" for a reason, so it should go against the normal (forward) direction
I think its fine, and I haven't heard others complain about it over the years.
Another way that this "upside down" way works for me is that this isn't water flowing downhill, it's being pushed up with every layer of the network adding energy or input.
Finally there's the metaphor of the roots of a plant being under the fruit of the plant.
(I didn't read the reddit post, apologies if it's a duplicate or these examples are addressed there.)
say root network
What is backpropagation really doing? https://youtu.be/Ilg3gGewQ5U
His other videos on this topic are just as good.
https://google-developers.appspot.com/machine-learning/crash...
Definitely worth checking out.
Only one, I hope constructive, criticism: Too many formulas without numbers. It will help the explanation if you include numbers and how the results are calculated. Not everybody is comfortable with the chain rule to distribute the error across the individual weights
It also skips over the bias value in the back propagation step.
He sort of goes through some other implications which, from an intuition massaging pov, is great.
https://www.youtube.com/watch?v=i94OvYb6noo
For those who are impressed by such things -- as I am -- he is now head of AI or ML or something at Tesla.
Just text accompanied by great visualisations.
Kudos!
> f(x) has to be a non-linear function, otherwise the neural network will only be able to learn linear models.
I thought one of the most common functions was relu, which is linear (but cuts off to 0 for x values below 0)
Any tool that combines these things is by necessity going to limit your creativity with them. Using the tools themselves is your best bet.
Check out https://mathisonian.github.io/idyll/scaffolding-interactives... (scroll example is towards the bottom)