Researchers propose faster, more efficient alternatives to backpropagation
venturebeat.com
venturebeat.com
> Another disadvantage of backpropagation is its tendency to become stuck in the local minima of the loss function. Mathematically, the goal in training a model is converging on the global minimum, the point in the loss function where the model has optimized its ability to make predictions.
"Backpropagation" is the method how to compute the gradient of the weights with respect to a loss function. But the article repeatedly uses the term as if it was the whole optimization algorithm, running into local minima.
venturebeat.com
Not the worst, but we're not talking Nature or Spectrum here.
-- Meaning follows usage.
Do you think a peer-reviewd publication is formal enough to warrant precise language? These publications are not only read by specialists in the field. I use automatic differentiation in my daily work, but I'm not familiar with machine learning. Thus I am very confused when "backpropagation" is used to mean an optimization algorithm.
EDIT: It is as if physicists used the term "special relativity" to talk about "quantum mechanics" because, after all, quantum mechanics happens in Lorentzian spacetime. Now for specialists of quantum physics it may make sense, since they are using "special relativity" to distinguish it from fancier quantum theories that combine field theory with GR. But for normal people it would be certainly misleading. Using "backpropagation" to include optimization has the same feeling.
Good day,
I was wondering whether it is be possible for you to provide an overview of different methods that you think might have a better shot at replacing backpropagation algorithm?
Reverse-mode differentiation has about the same time cost as whatever function you're optimizing, no matter how many parameters you need gradients for. This which is about as good as one could hope for, and is what lets it scale to billions of parameters.
The main downside of reverse-mode differentiation (and one of the reasons it's biologically implausible) is that it requires storing all the intermediate numbers that were computed when evaluating the function on the forward pass. So its memory cost grows with the complexity of the function being optimized.
So the main practical problem with reverse-mode differentiation + gradient descent is the memory requirement, and much of the research presented in the workshop is about ways to get around this. A few of the major approaches are:
1) Only storing a subset of the forward activations, to get noisier gradients at less memory cost. This is what the "Randomized Automatic Differentiation" paper does. You can also save memory and get exact gradients if you re-construct the activations as you need them (called checkpointing), but this is slower.
2) Only training one layer at a time. This is what the "Layer-wise Learning" papers are doing. I suppose you could also say that this is what the "feedback alignment" papers are doing.
3) If the function being optimized is a fixed-point computation (such as an optimization), you can compute its gradient without needing to store any activations by using the implicit function theorem. This is what my talk was about.
4) Some other forms of sensitivity analysis (not exactly the same as computing gradients) can be done by just letting a dynamical system run for a little while. Barak Pearlmutter has some work on how he thinks this is what happens in slow-wave sleep to make our brains less prone to seizures when we're awake.
I'm missing a lot of relevant work, and again I don't even know all the work that was presented at this one workshop. But I hope this helps.
Is it possible to combine these methods in a straight forward manner with methods that try to reduce the space complexity? For example, Lottery ticket hypothesis(https://arxiv.org/abs/1803.03635) seems to reduce spacial complexity(Please do correct me if I am wrong).
Also, based on my rather poor and limited knowledge, it appears to me that set of proposed methods that reduced space complexity and set of proposed methods that reduce time complexity are disjoint. Is that the case ?
There is a lot of work on trying to speed up optimization, for example the K-FAC optimizer by Roger Grosse that uses second-order gradient information in a scalable way.
The lottery ticket pruning strategies do reduce space complexity, but I think the main reason people are interested in it is to reduce training time complexity, or deployment memory requirements, but not so much training memory requirements.
As for whether memory-saving and time-saving approaches are disjoint, many methods (like checkpointing) introduce a tradeoff between time and space complexity, so no.
I wish you and your family a happy Christmas :)
Interesting! I am more familiar with Pearlmutter's work on automatic differentiation, but was was unaware of this work with Houghton.
A new hypothesis for sleep: tuning for criticality: https://zero.sci-hub.se/2153/6c1cfbc1b78d23ef2e1cb7102dd8339...
There is also a related paper on wake-sleep learning from UofT, of which I am sure you are aware:
The wake-sleep algorithm for unsupervised neural networks: https://www.cs.toronto.edu/~hinton/absps/ws.pdf
Are you aware of any recent work investigating the role of sleep in biological and statistical learning?
Basically if the network made the correct prediction tell each neuron to do a little bit more of what it just did. If it sent a high output, change the weights so it sends an even higher output. Weaken connections that were inhibitory and strengthen connections that were excitatory. And for a neuron with a low output, make it even lower by doing the opposite.
If on the other hand prediction was wrong, then try to make the neuron do less of what it did.
Do you know if something like this has been tried?
So my guess is that this approach would either take a much longer time to converge (as there's less information transmitted back for the neuron updates) or stall out completely.
Probably not too hard to code up, if you want to try it. But I would also be pretty surprised if it hadn't been tried before.
Slightly tangential question, but on reading the article I was surprised that it mentioned multiple anonymous submissions, given that this was a workshop at one of the most prestigious ML conferences. Any particular reason for this that you can think of?
Furthermore, the journalist managed to write an entire article about a workshop without once naming it, giving a link, or even defining backpropagation.
Edit: I want to say that I understand it's hard to cover an area that you're not an expert in. But if the journalist had googled the title of the articles he was writing about, he would have found the authors' names. Instead he gave the reader the impression that the articles are still anonymous.
I can easily write an article, "deep learning is dead, here comes quantum ml"
How you academic stop this bullshit, or are you yourself a bullshit who spread misinformation by supporting these kind of articles?
I don't really know what to do about the state of journalism, however. I thought I was being pretty hard on the author of this piece elsewhere in this thread, actually. I guess you could blame everyone who upvoted the article. On the other hand, we can't let the perfect be the enemy of the good.
There is a screening process for NIPS workshops (which I've been part of the last two years) where more experienced researchers rate the proposed workshops. There are usually a couple that get through that I think shouldn't be. However, I'd rather let 1000 stupid workshops through than shut down the one seemingly weird one that's actually a promising direction. Because if not here, where can we nurture and support weird ideas?
Finally, I think workshops play a vital role in onboarding newcomers to the field. One of the main ways you learn is by writing papers, and workshops lower the barrier to getting something out the door and getting feedback on it.
"If you want to learn flying by modeling the biology of birds, you're doing it wrong. Just look at today's airplanes. They have no resemblance to birds at all. Yet they're million times better and faster than any bird."
I hold a pessimistic view that we are still in hunter-gatherer mode as it comes to understanding cognition.
At some point you have to strike rocks to make fire, because the butane lighter hasn't been invented yet. You make do with what's available, and progressively get better at it. I tend to think that we're a couple-few perspective shifts away from getting it 'right,' and that the hardware side likely barely matters. But, I'm an optimist.
Having said that, you can certainly improve a design when you better understand the fundamentals (vs intuition + trial & error).
LMAO
A 6 years old kid can see the fundamental resemblance between a bird and a modern passenger airplane: The wings Tail stabilizer Slender body
Planes are faster bigger
Are they better?
Not necessarily, for example, humming bird can fly in a way that is far beyond any human machine in terms of efficiency and flexibility.
Of course man should not imitate birds, because human flight is fundamentally different activity than bird flying. But to say human aviation did not start by mimicking birds, is like to say Ann was not inspired human brain...
Birds are to planes, as humans are to cars. Yet can a car leap over barricades, climb mountains, trees, self-repair, turn on a dime, stop instantly, etc, etc?
A plane cannot maneuver like a bird, take off in crazy weather conditions, land on a dime in a tree, stop almost instantly in flight, and change direction, etc.
I think what you've quoted has a lot of value here, for, what we should expect from an artificial brain, isn't a human brain. This is truth. However, while it may be faster in a specific capacity, but it won't have the same characteristics.
So yes, expecting it to be like a human brain doesn't make sense.
Yet better/faster? I don't think we can compare this, they're too different.
(which is really the quote's point, but I just didn't like the better/faster bit at the end...)
The advanced models like GPT-3 are burning millions of watts in the cloud but they're not that much better than what a brain can do (and in many ways worse, as in often requiring supervised learning)
That's the key point. The algorithms need to become more energy efficient to make significant leaps, thus become more like brains.
This whole HN discussion of bird flight is a trainwreck and reflects massive gaps in understanding of aerodynamics. This is '00s "computer virus news report" level competence in this subject.
Actually, I'd say that our understanding of intelligence is right about at the level of aerodynamics at the dawn of heavier than air flight:
I mean, we could quibble about exactly where we are pre- or post-Wright Flyer, but given the amount of AI research that amounts to brute-force flailing about in search of incremental improvements, disagreements on the importance of "biological plausibility" and so on, it's pretty clear that, roughly speaking, AI is currently somewhere in the equivalent of the Lilienthal-Langley-Wright-Curtis continuum (ie. 1890-1910-ish) and still prior to the most important theoretical breakthroughs. IOW, AI has not in my opinion yet achieved an equivalent to aerodynamics' Prandtl lifting-line theory: https://en.m.wikipedia.org/wiki/Lifting-line_theory
There were similar arguments when AlphaGo showed up and beat master Go player Lee Sedol, but is power(in Watt) the right measurement? I always feel like it should be the total energy(in J or Cal) required to transform a computing device like biological brain or electrical computer from knowing nothing to being capable of a skill like Go game. In such sense, deep learning is still more energy efficient than human.
Better/faster we would not directly compare to humans, but to benchmarks and timed experiments.
LeCun is saying to treat "intelligence" the same as "flight" or "swimming". It is a matter of function, not a matter of a specific instantiation on a biological substrate. You don't need to recreate flapping wings to gain "flight", you can strap a combustion engine on a cylinder and beat all birds on earth in regards to speed. You don't say "we don't have flight yet", because an airplane is not able to land on a tree branch. Maybe we don't have yet all the components and aspects of "flight", but this is not a show stopper, and drones have come a long way.
Now the more interesting question becomes: What are the laws of aerodynamics for intelligence?
Aside: I think it is absolutely insane that a conference workshop with papers yet to go through peer-review, is highlighted as a popsci article on VentureBeat. That's such a narrow workshop, that even researchers in the field may be unaware of it. And now these get to read the paper summaries from a HN-story. "the centre cannot hold".
Aside II: Yann LeCun talk from 2019 about this subject (better to debate the source ;)):
> Clearly, Deep Learning research would greatly benefit from better theoretical understanding. DL is partly engineering science in which we create new artifacts through theoretical insight, intuition, biological inspiration, and empirical exploration. But understanding DL is a kind of "physical science" in which the general properties of this artifact is to be understood. The history of science and technology is replete with examples where the technological artifact preceded (not followed) the theoretical understanding: the theory of optics followed the invention of the lens, thermodynamics followed the steam engine, aerodynamics largely followed the airplane, information theory followed radio communication, and computer science followed the programmable calculator. My two main points are that (1) empiricism is a perfectly legitimate method of investigation, albeit an inefficient one, and (2) our challenge is to develop the equivalent of thermodynamics for learning and intelligence. While a theoretical underpinning, even if only conceptual, would greatly accelerate progress, one must be conscious of the limited practical implications of general theories. --- https://www.ias.edu/video/DeepLearningConf/2019-0222-YannLeC...
See my link to his ICML 2013 presentation above.
We don't have the same kind of understanding of how brains learn, so the comparison is not quite right.
When we understand how to build things that learn like brains, we'll be in a better position to say things like "Ok this is strictly worse than backprop, let's stick with backprop" or "Actually, this is better than backprop because X", (or, more likely, there are things we can use from both). Until we have that understanding it's silly to stop trying to understand how the brain does things.
That being said, nobody is going to stop working on backprop, and no one is going to stop working on understanding biological mechanisms . Research works by a bunch of people investigating different avenues simultaneously.
> "The question of whether machines can think is about as relevant as the question of whether submarines can swim."
https://cilvr.nyu.edu/lib/exe/fetch.php?media=deeplearning:2...
Slide #9
Let's be inspired by nature, but not too much
It's nice imitate Nature,
But we also need to understand
For airplanes, we developed aerodynamics and compressible fluid dynamics.
Question : what is the equivalent of aerodynamics for understanding intelligence
But I like the gist of the quote.
This is the NeurIPS workshop that the article is talking about.
"But how do we select a good network from these Kn different networks? Brute-force evaluation of all possible configurations is clearly not feasible due to the massive number of different hypotheses. Instead, we present an algorithm, shown in Figure 1, that iteratively searches the best combination of connection values for the entire network by optimizing the given loss. To do this, the method learns a real-valued quality score for each weight option. These scores are used to select the weight value of each connection during the forward pass. The scores are then updated in the backward pass based on the loss value in order to improve training performance over iterations."
It's actually pretty clever.
However, it's still quite hard to get useful results from it in practice.
---------------
I personally believe that one component of intelligence is the ability to apply cognitive patterns created for a particular input to other inputs. (Very simplified example: A block of "neurons" which have learned to recognize the pattern "is hurt by" when given a subject (group of pixels in image) and object (other group of pixels in image), could be applied to another subject/object pair, for example coming from processed audio. But if the audio processing takes 10 layers, and the image processing 5, the connection has to run backwards)
To do this in a state of the art deep network, you need the ability to create backward connections. Backward connections imply loops, and loops break backprop (unlike loops in RNNs, which can be easily unrolled AFAIK). So with the current backprop-trained feedforward model, you have to create patterns multiple time instead of reusing them.
This is why I will pay attention to backprop alternatives which allow loops, despite their (currently many) disadvantages. This and modular training are the two aspects of learning I would personally focus on.
I am not exactly working on this, however, intuitively believe that it should be possible and effective to make STDP work (including handling loops).
Thanks.
SLIDE seems way, way superior to any of the listed solutions or approaches, as far as I could tell on a first read through.
I'm not sure if its related, but would this work kind of how armadillo can do singular value decomp [0] of a matrix by embedding arbitrary n by m matrix X in a higher dimensional n+m by n+m null matrix M?
The gradient in machine learning is based on the loss. Specifically it's the direction that reduces the loss the fastest. So, not only the most recent batches, but specifically by the recent data that is predicted incorrectly. It doesn't have any "confidence" from the memory of what was predicted right previously, for example, it just currently only cares about changing to suit the most recent batches.
It's effective, but the so-called "boundary cases" often have to be hand-chosen due to the difficulty of selecting them automatically: Early samples always have high loss and the decision boundary nearness is implicitly connected to the accuracy of the network at time of evaluation. In other words, the function we evaluate on forward pass itself is changing as a result of backprop, so the critical points and output of the function are also in flux.
You also lose an increasing portion of each batch to the "important" cases as you add more, so maintaining the size and contents of this pool is difficult - if you added every case, you'd have no new data.
So I think it's promising, but it needs more foundational work on deriving the impact of individual samples on the output. (If we ever get that breakthrough in explainability...)
Overall I tend to think that this space is underexplored compared to searching for new architectures... We know that it helps to choose a curriculum for humans to help guide learning, even beginning with 'baby talk' to develop early communication skills.
Are anonymously submitted papers becoming (more) common? If so, what's driving this?
Ah, that makes sense.
Hah!
So, invention recapitulates evolution?
Feedback for the purpose of regulating the state of a machine in response to input dates to antiquity, if we're really getting absurd. The formal definition is also debatable, I think Maxwell has the strongest claim.
It was clearly phrased this way specifically because backprop is just the chain rule applied in a particular direction, and as such has been invented and reinvented over and over by every one under the sun. Hell, a lazy googling says gradient descent goes back to Cauchy.