Anyway, residual connections in NNs as well as distillation being only a 1% hit to performance imply our models are way too big.
Anyway, residual connections in NNs as well as distillation being only a 1% hit to performance imply our models are way too big.
I disagree with the conclusion.
It indicates that our optimisers are just not good enough, likely because gradient descent is just weak.
The argument for residual connections is that we can create a nested family of models which enables expressing more models, but also embedding the smaller ones into them.
The smaller models may be retrieved if our model learns to produce the the identity function at later layers.
The problem though is that that is very difficult, meaning that our optimisers are simply not good enough at constructing identify functions. With the residual layers, we can embed the identity function into the structure of the model, and we now need to learn to map to 0 (since a residual is f(x) = x+g(x)), we need only to learn g(x)=0).
As for our optimisers being bad, the argument is that with an overparameterised network, there is always a descent direction, but we land on local minima that are very close to the global one. The descend direction may exist in the batch, but when considering all the batches, we are at a local minimum.
We can find many such local minima via certain symmetries.
The general problem however is that even with the full dataset, we can only make local improvements in the landscape.
Thus, it’s that the better models are embedded within the larger ones, and more parameters enable us to find them because of nested families, symmetries, and because of always having a descent direction.
No, the networks are ok, what is wrong is the paradigm. If you want rule or code based exploration and learning it is possible. You need to train a model to generate code from text instructions, then fine-tune it with RL on problem solving. The code generated by the model is interpretable and generalises better than running computation in the network itself.
Neural nets can also generate problems, tests and evaluations of the test outputs. They can make a data generation loop. As an analogy, AlphaGo generated its own training data by self play and had very strong skills.
I don’t think that this reply takes into consideration just how inefficient RL and the likes are. In fact, RL is so inefficient that current SOTA in RL is … causal transformers that perform in-context learning without gradient updates.
Depending on the approach one takes with RL, be it policy gradients or value networks, it still relies on gradient descent (and backprop).
Policy gradients are just increasing the likelihood of useful actions given the current state. It’s a likelihood model increasing probabilities based on observed random walks.
Value networks are even worse because one needs to derive not only the quality of the behaviour but also select an action.
Sure enough, alternative methods exist such as model based RL, etc, and for example ChatGPT use RL to train some value functions and learn how to rank options, but all of these rely on gradient descent.
Gradient descent, especially stochastic, is just garbage compared to stuff that we have for fixed functions that are not very expensive to evaluate.
With stochastic gradient descent, your loss landscape depends on the example or mini batch, so a way to think about it is that the landscape is a linear combination of all the training examples, but at any time you observe only some of them and cope that the gradient doesn’t mess up too bad.
But in general gradient descent shows linear convergence rate (cf Nocedal et al Numerical Opt, or Boyd and Vanderberghe’s proof where they bound the improvement of the iterates), and that’s a best case scenario (meaning non stochastic, non partial).
Second order methods can get quadratic convergence rate but they are prohibitly expensive for large models, or require hessians (good luck lol).
None of these though address limitations imposed by loss functions, eg needing exponentially higher values to increase a prediction optimised by cross entropy (see the logarithm). Nor do they address the bound on the information that we have about the minima.
So needing exponentially more steps (assuming each update is fixed in length) while relying on linear convergence is … problematic to say the list
To generate more, you need to interact with an environment, and you also need an objective function. If we could magically generate more textual data, then we have a language model already and don't need to train another language model.
You can't bootstrap a model for language synthesis unless you give it access to the internet to interact with users, at which point ... you have a Tay [1]
If you have any ideas to do better and aren't idly wealthy, I'd suggest pursuing them. Create a model that's within a percentage point or two of GPT3 on big NLP benchmarks, and fame and fortune will be yours.
[Edit] this of course only applies for domains like NLP or computer vision where neural networks have proven very hard to beat. If you're working on a problem that doesn't need deep learning to achieve adequate performance, don't use them!
My perspective is more pessimistic. I think people opt for huge unsupervised models because they believe that tuning a few thousand more input features is easier than labeling copious amounts of data. Plus (in my experience) supervised models often require a more involved understanding of the math, whereas there's so many NN frameworks that ask very little of the users.
Companies like Google have even spent huge amounts of time and money on enormous labeled datasets -- JFT-300M or something like that for computer vision tasks, as you might guess, ~300M labeled images. It creates value, but it creates more value for larger models with higher capacity.
And yet people still like to push this idea that we will magically and accidentally build a superintelligence on top of these systems. It's so frustrating how deep into their own koolaid the ML industry is. We don't even know how the brain learns, we don't understand intelligence, there's no valid reason to believe a NN "learns" the same way a human brain learns, and individual human neurons are infinitely more complex and "learning" than even a single layer of a NN.
No, we understand very well how NNs work. Look at PartiallyTyped's comment in this thread. It's a great explanation of the basic concepts behind modern machine learning.
You're quite correct that modern neural networks have nothing to do with how the brain learns or with any kind of superintelligence. And people know this. But these technologies have valuable practical applications. They're good at what they were made to do.