From what I understand, gradient descent and its cousins can't suddenly jump to a distant global optimum.
From what I understand, gradient descent and its cousins can't suddenly jump to a distant global optimum.
It takes 3 minutes to train the Shakespeare model with gradient descent. The black-box methods I tested so far likely take 30+ hours to train (I haven't tried to take them to the end yet). I've hit a wall where progress is very slow. The text generated at that stage has punctuation and words are split with spaces but the words themselves are mostly nonsense. Almost feels like it learned that English is letters separated by spaces, and that you put exclamation marks or periods at the end but not that much more.
There's some larger scale CMA-ES variants I still want to test that don't have quality implementations. I've tried to stare at pictures of gradients and weights from half-trained models and trying to come up with ideas how to get there with black-box optimization. Also trying some original ideas where you compute a gradient, but you would not compute it against a loss function. The gradient would be more for discovering hidden structure in weights, that you would then put on some black-box optimizer as a guide (which I guess makes it not entirely black box. Gray box?)
Possible? I mean, I guess technically. Practical? No way, unless some major breakthrough happens.
My current goal is to just produce a model, even if training takes laughably long so I can say I've trained a language model using nothing but getting a fitness score from a black box function.
Edit: if you are reading this and are aware of any other serious attempts at training a non-trivial sized language model without gradient descent I would want to know. So I can try their methods. I know there's some large scale stuff used in reinforcement learning like in one Uber paper but not in LLMs specifically.
It's about information. Gradient-free methods integrate little or no information about the problem; they're a blind watchmaker. This works, but it's slow and gets slower the bigger your problem is. (curse of dimensionality)
Gradients integrate some limited information about the problem. This lets you find solutions much faster, and neural networks are structured specifically to be easy to optimize with gradients. Local minima don't seem to be a problem.
The future is probably even smarter optimizers that integrate more information about the problem and learn to make good assumptions. This is the goal of Learned Optimizers, like Velo (https://arxiv.org/abs/2211.09760).
But it doesn't solve the problem of local minima, and it will also need to use minibatches.
https://www.publichealth.columbia.edu/research/population-he...
https://en.wikipedia.org/wiki/Kriging
Basically, you are wasting most of your compute to come up with a rough local approximation to the thing you actually want. But that's sort of pointless in the NN training context, because what you want is basically the gradient (and maybe some higher order terms that tell you about the local curvature too).
CMAES makes sense when the gradient is not even well defined. For example, if you have a bunch of parameters for an airplane design, and then want to take that design and do a bunch of huge aerodynamics calculations to compute its lift, or do a big finite element analysis to measure how well it withstands various stresses, and at the end of that big analysis, you get back a number, like "maximum lift" or something. If each run takes hours on a supercomputer, then you clearly don't have anything close to a gradient and it would be very expensive to even try to approximate it numerically. So CMAES is useful there in helping you pick better high level parameters in a smart way-- basically it's a big improvement over grid search.
I haven't estimated the number of trials you would need for 10M Shakespeare model but I think to get to the same level as gradient descent, it might be around 10M, i.e. same ballpark as the number of parameters. Which makes some intuitive sense because of how little you learn from each black box function evaluation.
There's maybe some hope that there is a hidden structure in the problem that does not actually need anywhere close to 10M parameters so that a black box optimizer might still find it. I don't have my hopes up though but I'm trying to poke at it.
I would think that if it turns out LLMs are not totally impossible with black box optimization, then it would be good to find a reason to use it. E.g. objective functions that don't have a hope of having a good gradient. Some kind of weird architecture that can't be trained conventionally. Maybe fine-tuning gradient descent optimized models with those things. Etc. Feels like a solution looking for a problem.
I'm doing my project for fun and just seeing if it's possible at all.
Next up after this project is that I want to test some metalearning ideas. I read some papers where the idea is that all weights are actually tinier neural networks, all with the same parameters where you train it to learn backpropagation (or whatever learning algorithm it converges to). The paper I read this from argued it also worked for forward-only but my intuition doesn't quite understand how. I want to follow up a bit on this line of research and check if there's been any new developments since I read them and try them out in my own code.
When your parameter space is in the order of billions, for all practical purposes, there is always a direction of descent.
More over, local minima seem to be rather close to the global minima.
https://www.google.com/url?sa=t&source=web&rct=j&opi=8997844...
(Edit: eh sort of ...)
Here's me implementing an algorithm from 2009 in single-core on a CPU and getting pretty excellent results on RLHF benchmarks: https://github.com/Hellisotherpeople/Python-Cooperative-Syna...
Don’t many/most state of the art models take many months to train on far more data than humans need for similar tasks?
Also, while e.g. GPT4 is quite capable across many tasks - humans seem to average towards learning robust _learning techniques_ themselves. Learning a new subject becomes easier thanks to somehow tracking and encoding learning strategies that are robust to learning other unrelated topics.
Humans generally need 18 years of pre training followed by 4-6 years of fine tuning before they can “one-shot” many difficult tasks. That’s way more training than any machine learning model I’m aware of.
Even for tasks like reading the newspaper and summarizing what you read, you probably had to train for 10-12 years.
A 3 year old has 3 years of multimodal training data and RLHF + a few billions of years of evolution that have primed and biased our visual and cognitive systems.
That requires a lot more data than machine models that literally zero inherent bias. Assuming you want a true apples to apples comparison.
If you want to be pedantic then only 6% of the human brain is the visual cortex. But AlexNet is also an inefficient model so something like an optimized ResNet is 100x as efficient to train. So now you're at 10.5kwh and 1.5kwh for the baby and model respectively.
You can argue details further but I'd say the energy cost of both is fairly close.
The networks we have are trained once and work efficiently for their training dataset. They are even robust to outliers in the distribution of that dataset. But they aren’t robust to surprises, unmentioned changes in assumptions/rules/patterns about the data.
Even reinforcement learning is still struggling with this as self-play effectively requires being able to run your dataset/simulation quickly enough to find new policies that work. Humans don’t even have time for that, much less access to a full running simulation of their environments. Although we do generate world models and there’s some work towards that I believe.
Again happy to be corrected.
If you want to be pedantic then 6% of the human brain is the visual cortex but then you also have to argue that AlexNet is horribly inefficient to train. So you cut the brain cost to 6% and the model cost to 1%. They're still within an order of magnitude (favoring the model) which I'd say is pretty close in terms of energy usage.
Humans are very bad at this task; it takes a massive effort to learn this many birds. In fact it's a great counterexample to human few shot learning ability...