I can easily write an article, "deep learning is dead, here comes quantum ml"
How you academic stop this bullshit, or are you yourself a bullshit who spread misinformation by supporting these kind of articles?
I don't really know what to do about the state of journalism, however. I thought I was being pretty hard on the author of this piece elsewhere in this thread, actually. I guess you could blame everyone who upvoted the article. On the other hand, we can't let the perfect be the enemy of the good.
There is a screening process for NIPS workshops (which I've been part of the last two years) where more experienced researchers rate the proposed workshops. There are usually a couple that get through that I think shouldn't be. However, I'd rather let 1000 stupid workshops through than shut down the one seemingly weird one that's actually a promising direction. Because if not here, where can we nurture and support weird ideas?
Finally, I think workshops play a vital role in onboarding newcomers to the field. One of the main ways you learn is by writing papers, and workshops lower the barrier to getting something out the door and getting feedback on it.
Good day,
I was wondering whether it is be possible for you to provide an overview of different methods that you think might have a better shot at replacing backpropagation algorithm?
Reverse-mode differentiation has about the same time cost as whatever function you're optimizing, no matter how many parameters you need gradients for. This which is about as good as one could hope for, and is what lets it scale to billions of parameters.
The main downside of reverse-mode differentiation (and one of the reasons it's biologically implausible) is that it requires storing all the intermediate numbers that were computed when evaluating the function on the forward pass. So its memory cost grows with the complexity of the function being optimized.
So the main practical problem with reverse-mode differentiation + gradient descent is the memory requirement, and much of the research presented in the workshop is about ways to get around this. A few of the major approaches are:
1) Only storing a subset of the forward activations, to get noisier gradients at less memory cost. This is what the "Randomized Automatic Differentiation" paper does. You can also save memory and get exact gradients if you re-construct the activations as you need them (called checkpointing), but this is slower.
2) Only training one layer at a time. This is what the "Layer-wise Learning" papers are doing. I suppose you could also say that this is what the "feedback alignment" papers are doing.
3) If the function being optimized is a fixed-point computation (such as an optimization), you can compute its gradient without needing to store any activations by using the implicit function theorem. This is what my talk was about.
4) Some other forms of sensitivity analysis (not exactly the same as computing gradients) can be done by just letting a dynamical system run for a little while. Barak Pearlmutter has some work on how he thinks this is what happens in slow-wave sleep to make our brains less prone to seizures when we're awake.
I'm missing a lot of relevant work, and again I don't even know all the work that was presented at this one workshop. But I hope this helps.
Is it possible to combine these methods in a straight forward manner with methods that try to reduce the space complexity? For example, Lottery ticket hypothesis(https://arxiv.org/abs/1803.03635) seems to reduce spacial complexity(Please do correct me if I am wrong).
Also, based on my rather poor and limited knowledge, it appears to me that set of proposed methods that reduced space complexity and set of proposed methods that reduce time complexity are disjoint. Is that the case ?
There is a lot of work on trying to speed up optimization, for example the K-FAC optimizer by Roger Grosse that uses second-order gradient information in a scalable way.
The lottery ticket pruning strategies do reduce space complexity, but I think the main reason people are interested in it is to reduce training time complexity, or deployment memory requirements, but not so much training memory requirements.
As for whether memory-saving and time-saving approaches are disjoint, many methods (like checkpointing) introduce a tradeoff between time and space complexity, so no.
I wish you and your family a happy Christmas :)
Interesting! I am more familiar with Pearlmutter's work on automatic differentiation, but was was unaware of this work with Houghton.
A new hypothesis for sleep: tuning for criticality: https://zero.sci-hub.se/2153/6c1cfbc1b78d23ef2e1cb7102dd8339...
There is also a related paper on wake-sleep learning from UofT, of which I am sure you are aware:
The wake-sleep algorithm for unsupervised neural networks: https://www.cs.toronto.edu/~hinton/absps/ws.pdf
Are you aware of any recent work investigating the role of sleep in biological and statistical learning?
Basically if the network made the correct prediction tell each neuron to do a little bit more of what it just did. If it sent a high output, change the weights so it sends an even higher output. Weaken connections that were inhibitory and strengthen connections that were excitatory. And for a neuron with a low output, make it even lower by doing the opposite.
If on the other hand prediction was wrong, then try to make the neuron do less of what it did.
Do you know if something like this has been tried?
So my guess is that this approach would either take a much longer time to converge (as there's less information transmitted back for the neuron updates) or stall out completely.
Probably not too hard to code up, if you want to try it. But I would also be pretty surprised if it hadn't been tried before.
Slightly tangential question, but on reading the article I was surprised that it mentioned multiple anonymous submissions, given that this was a workshop at one of the most prestigious ML conferences. Any particular reason for this that you can think of?
Furthermore, the journalist managed to write an entire article about a workshop without once naming it, giving a link, or even defining backpropagation.
Edit: I want to say that I understand it's hard to cover an area that you're not an expert in. But if the journalist had googled the title of the articles he was writing about, he would have found the authors' names. Instead he gave the reader the impression that the articles are still anonymous.