Geoffrey Hinton publishes new deep learning algorithm
infoq.com
infoq.com
Those are relatively close figures, but good accuracy on CIFAR-10 is 99%+ and getting ~94% is trivial.
So, if an improper architecture for a problem is used and the accuracy is poor, how compelling is using another optimization approach and achieving similar accuracy?
It's a unique and interesting approach, but the article specifically mentions it gets accuracy similar to backprop, but if this is the experiment that claim is based on, it loses some credibility in my eyes.
Almost any ML algorithm can be thrown at CIFAR10 and achieve ~60% accuracy; this ballpark of accuracy is really not sufficient to demonstrate viability, no matter how aesthetically interesting the approach might feel.
But "any ml algorithm" isn't the point. It's a new optimization technique and should be applied to models/architectures that make sense with the problems they are being used on.
For example, they could have used a pretrained featurizer and trained the two layer model on top of it, with both back prop and FF and compared.
Making the assumption that weights/embeddings produced by a backprop-trained network are equally intelligible to a network also trained by backprop vs. one trained by this alternative method.
That said, I am also short-term bearish on backprop-free methods (although potentially long-term bullish).
If he invents the new back propagation, an army of grad students can turn his ideas into the future. Like they've done for the last 15 years.
He's posting incremental work towards rethinking the field. It's pretty interesting stuff.
Edit: grammar
If, however, you are ripping out backpropagation like this paper is, then you get a big pass. This is not the new paradigm yet, but it's promising that it doesn't just completely fail.
It took >20 years for it to be right.
Maybe we ought to give this one some time?
The failure mode of tenure is that the professor just rests on their past accomplishments and doesn't do anything. That's a risk the system takes. In this case though, Geoff Hinton is doing everything right: he's not only not sitting around doing nothing, he's actively trying to obsolete the paradigm he helped usher in, just in case there is a better option out there. I think that's admirable
Until the brain's algorithm is "solved", half steps are important. We need as many alternate half steps as we can find until one or more lead to a better understanding of the brain. (And potentially, better than backdrop efficiency or results.)
But, model size and complexity also matters. MLP with backprop gets about 63% on cifar-10, for various reasons. So achieving 59% accuracy means this algorithm is about 93% as good as backprop in this case.
However, 63% accuracy on cifar-10 can be achieved with two (maybe three) layers IIRC. The output is a 10-way classifier, which is handled in one layer. If the output requires multi-layer transformations, then gradients need to be back-propagated.
As long as the batch activation vectors are trained to max separation (or orthogonality or whatever) at each layer, one output layer can match them to labels. But this is unlikely in problems where the output is more "transformed or complicated".
So how do you optimize a layer? Do you still use gradient descent? So you are have a per layer loss with a positive and negative component and then do gradient descent?
So then what is the label for each layer? Do you use the same label for each layer?
And what does he mean by the forward pass not being fully known? I don't get this application of the blackbox between layers. Why would you want to do that?
Not accurate for the version another commenter linked: https://www.cs.toronto.edu/~hinton/FFA13.pdf
I see four equations.
The idea is imo similar to how random word embeddings are generated.
I'm not sure where you get that impression. Forward-Forward [1] seems to eschew gradients entirely:
The Forward-Forward algorithm replaces the forward and backward passes of backpropagation by two forward passes, one with positive (i.e. real) data and the other with negative data which could be generated by the network itself
[1] https://www.cs.toronto.edu/~hinton/FFA13.pdfThe implementations of this compute gradients locally.
In linear layers, it is possible. Once you have computed the gradient of the output of the vector ith vector, so a scalar, you scale the input by that value and add it to the parameters.
This is a simple FMA op: a=fma(eta*z, x, a), with z the gradient of the vector, x the input, a the parameters, and eta the learning rate. This computes a = a + eta*z*x in place.
Unfortunately I've also seen papers get rejected because their idea was "trivial", yet no one had thought of it before. Hinton has an edge here though.
Scaling Forward Gradient With Local Losses Mengye Ren, Simon Kornblith, Renjie Liao, Geoffrey Hinton https://arxiv.org/abs/2210.03310
and has code: https://github.com/google-research/google-research/tree/mast...
I don't get it, don't all of those optimizers work via backprop?
> 7 The relevance of FF to analog hardware
> An energy efficient way to multiply an activity vector by a weight matrix is to implement activities as voltages and weights as conductances. Their products, per unit time, are charges which add themselves. This seems a lot more sensible than driving transistors at high power to model the individual bits in the digital representation of a number and then performing O(n^2) single bit operations to multiply two n-bit numbers together. Unfortunately, it is difficult to implement the backpropagation procedure in an equally efficient way, so people have resorted to using A-to-D converters and digital computations for computing gradients (Kendall et al., 2020). The use of two forward passes instead of a forward and a backward pass should make these A-to-D converters unnecessary.
It was my impression that it is difficult to properly isolate an electronic system to use voltages in this way (hence computers sort of "cut" voltages into bits 0/1 using a step function).
Have these limitations been overcome or do they not matter as much, as neural networks can work with more fuzzy data?
Interesting to imagine such a processor though.
https://www.nature.com/articles/s41467-020-20719-7
https://opg.optica.org/optica/fulltext.cfm?uri=optica-5-7-86...
If these FF networks can be proven to scale or made to scale similarly to BP networks, this would enable making hardware several orders of magnitude more efficient, for the price of loosing the ability to make exact copies of models to other computers. (The loss of reproducibility sits well with the tradition of scientific papers anyway/s;)
2.) How does this paper relate to Hintons feedback alignment from 5 years ago? I remember it was feedback without derivatives. What are the key new ideas? To adjust the output of each individual layer to be big for positive cases and small for negative cases without any feedback? Have these approaches been combined?
Nvidia isn't creating new versions of its NVLink/NVSwitch products just for the sake of it, better communication must be a key enabler.
Can someone with deeper knowledge can comment on this? Is communication a bottleneck, and will this algorithm uncover a new design space for NNs?
Communication across GPUs doesn’t solve this but instead allows to have either many models running in parallel on different GPUs to increase Barch size or to share many layers across GPUs to increase model size. Quick communication is critical to maintain training speeds that aren’t astronomical
No.
Hinton "discovered" stacking ensembles and gave it a new name, fancy analogies to biological brains and then made it worse.
The gist of this is that you can select a computational unit, be it a linear layer, or a collection of layers, compute the derivative of the output with respect to the parameters, and update them.
Each computational unit is independent, meaning that you don't calculate gradients going outside of it.
This is the same as training a bunch of networks, computing predictions, and then using another layer to combine the predictions. This is called stacking, and the networks are called an "ensemble". You can do this multiple times and have N levels of meta estimators.
Instead of fitting the ensemble and then the meta estimator, Hinton proposes training both simultaneously but without allowing gradients to flow through.
That is stupid because if you don't allow gradients to flow through, you will see a context drift as the data distribution changes. Hinton observed this context drift, to deal with that, he proposed normalizing the data.
On one extreme, you can use individual linear units as the models, and on the other extreme, you can combine all units into a single neural network and treat that as a module.
So no, this does not open any new design, it's an old idea, worsened, and wrapped in fancy words and post-facto reasoning.
If you are curious how a linear layer is an ensemble, observe that each vector is its own linear estimator, making the linear mapping an ensemble of estimators.
"Context drift as data distribution changes" sounds a hell of a lot like real life to me.
Normalized = hedonic treadmill on long view
At the micro scale, data that overflows the normalization is stored in emotional state, creating an orthogonal source of truth that makes up for the lack of full connected learning.
The fact that I argued why I found it bogus based on well established principles, and I get shitted on by people who by all means have provided nothing to this conversation and except suppressing criticism or throwing ad-hominems should tell all about the quality of discourse.
Dismissing criticism, not by arguments, but by the mere name of the person does a disservice to everyone.
If the research can't stand on its own, independent of the author, then it is not good research.
If you can point to _fundamental_ criticism of my arguments, and not fallacies or attacks, I'd be more than happy to discuss them.
With that kind of tonal promise, especially considering the source you are dismissing outright is important in their field, you have to show, not just tell.
If you just left that No out, and gave room for the chance that you are wrong, people wouldn't downvote, they'd upvote. People like to hear smart arguments. No one wants to hear outward dismissal. Especially of known experts.
I could name many other people who have actually been more influential in the field.
Since gradients don't flow from B to A in (B.f.A)(x), A is trained independently of B, meaning that the training distribution of B changes without B influencing it, i.e. context drift. B doesn't know the difference, and B doesn't influence it.
For all intents and purposes, you can compute all the outputs as training of A happens, meaning training A to completion, and then feed them into B and B will still compute the same outputs and derivatives as it did before.
To deal with context drift, Hinton proposes normalizing the data, so the distribution does not change significantly.
Whatever he proposed is not "backprop-free" either. It still involves backprop, but the number of layers gradients flow through is 1, the layer itself.
The argument that you can still train through non-differentiable operations is not particularly convincing either; the reparameterization trick shows that is trivial to pass gradients through non differentiable operations if we are smart about it.
Given non differentiable operator Z: R^N -> R^N; let A, B, C be R^N -> R^N linear layer, B(Z(C(x)) * A(C(x))) allows gradients to flow through B and A all the way to C. The output of Z is for all intents and purposes a Hadamard product with (A . C)(x) that is runtime constructed and might as well be part of the input.
You can even run Z(C(x)) through a neural network and learn how to transform that and still provide useful and informative gradients back to C(x) via (A . C)
By looking at it through that perspective, the issues with the approach become evident, and are fundamental in my opinion.
We extract learning, while we are imbibing the data and there seems to be no mechanism in the brain that favors backprop like learning process.
What Hinton and Deepmind will do is use neural-network learned-data, or perhaps the weights, as input to this kind of network. In other words, the output of another NN is labeled a priori, ergo you can use it "unsupervised" networks, which this research expounds. This will allow them to cook the input network into a specific dish, by labels even. Now give me my phd.
edit: edit
He also wrote an interesting thread on the memory usage of this algo versus backprop https://twitter.com/diegofiori_/status/1605242573311709184?s...
--
The divulgational title is almost an understatement: the Forward-Forward algorithm is an alternative to backpropagation.
Edit: sorry, the previous formulation of the above in this post, relative to the advantages, was due to a misreading. Hinton writes:
> The Forward-Forward algorithm (FF) is comparable in speed to backpropagation but has the advantage that it can be used when the precise details of the forward computation are unknown. It also has the advantage that it can learn while pipelining sequential data through a neural network without ever storing the neural activities or stopping to propagate error derivatives....The two areas in which the forward-forward algorithm may be superior to backpropagation are as a model of learning in cortex and as a way of making use of very low-power analog hardware without resorting to reinforcement learning
Meta question on HN implementation: Why do sometimes submitting a previously submitted resource links automatically to the previous discussion, while other times is considered a new submission?
https://hn.algolia.com/?query=Geoffrey%20Hinton%20publishes%...
This makes me cheerful because it suggests a way that studying systems which appear intelligent might be able to teach us more about how human intelligence works.
The confusing thing with this claim is what did people actually do during this time, given bad (and expensive!) lighting only?
> My guess, this sleep pattern is better for learning.
That might be true. One of the techniques to induce lucid dreaming works similarly – sleep for 4-5 hours, wake up, stay awake for 15-60mins then go back to sleep. It's called "wake back to bed" technique. Many lucid dreamers report increased capacity for learning in dreams.
Most of the "training process" of our brain likely occurred prior to our birth in evolutionarily optimized structure of brain.
We constantly find out that certain things are actually really important even though we thought it was junk. Recall that our best ability to test Genes is by knocking them out one by one and trying to observe the effect
The brain is comprised of many extremely specialized sub systems and formulas for generating knowledge. We don’t know English at birth, sure, but we do have a language processing capability. The training baked into the brain is a level of abstraction higher, establishing frameworks to learn other things. It may not be as storage data heavy, but it’s much harder to arrive at and is the bulk of the learning process (learning to learn)
As far as I can tell, this is almost the same as stacking multiple layers of ensembles, except worse as each ensemble is trained while previous ensembles are learning. This is causing context drift.
To deal with the context drift, Hinton proposes to normalise the output.
This isn't anything new or novel. Expressing "ThIs LoOkS sImIlAr To HoW cOgNiTiOn WoRkS" to make it sound impressive doesn't make it impressive or good by any stretch of the imagination.
Hinton just took something that existed for a long time, made it worse, gave it a different name and wrapped it in a paper under his name.
With every paper I am more convinced that the Laureates don't deserve the award.
Sorry, this "paper" smells from a mile away, and the fact that it is upvoted as much shows that people will upvote anything if they see a pretty name attached.
Edit:
Due to the apparent controversy of my criticism, I can't respond with a reply, so here is my response to the comment below asking what exactly makes this worse.
> As far as I can tell, this is almost the same as stacking multiple layers of ensembles
It isn't new. Ensembling is used and has been used for a long time. All kaggle competitions are won through ensembles and even ensembles of ensembles. It is a well studied field.
> except worse as each ensemble is trained while previous ensembles are learning.
Ensembles exhibit certain properties, but only iff they are trained independently from each other. This is well studied, you can read more about it in Bishop's Pattern recognition book.
> This is causing context drift.
Context drift occurs when a distribution changes over time. This changes the loss landscape which means the global minima change / move.
> To deal with the context drift, Hinton proposes to normalise the output.
So not only is what Hinton built a variation of something that existed already, made it worse by training the models simultaneously, and to handle the fact that it is worse, he adds additional computations to deal with said issue.
If it's not new or novel, why aren't people using it? If it's bad, what's wrong with it?
No comparisons to AdamW were made.
In fact, this algorithm uses backprop at its core, but propagating through 0 layers.
I don't get what he means by inserting the label into the input and what labels he is using per layer.
Obviously, this method is problematic if you have thousands of labels or if your network is not a classifier.