Researchers Build AI That Builds AI
quantamagazine.org
quantamagazine.org
For example, if we give the trained GNN a ResNet-50 architecture, the GNN predicts the trained weights in one forward pass, and voilà, we have a ResNet-50 that is now ready for use -- and you can always finetune it to get better performance out of it.
Remarkably, the prediction of weights works for previously unseen model architectures too. Design your own model, run its architecture by the GNN to predict the trained weights, and voilà, you have a model that is now ready for use -- and you can always finetune it to get better performance out of it.
Very cool!
--
[a] https://arxiv.org/pdf/2110.13100.pdf -- Abstract: "Deep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures – DEEPNETS-1M– and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis."
It's kind of like a more advanced version of hyper-parameter tuning which has been done in more primitive forms for a long time. But this approach definitely seems powerful.
It's for instantly training/pretraining any architecture for one or more predefined tasks (e.g., you could have a GNN for visual tasks, a GNN for NLP tasks, etc.), but you could also have a multi-domain GNN which accepts two inputs: an architecture and a domain.
Research use: Can we figure out how and why these GNNs can predict trained weights for previously unseen architectures without actually having to train the models? Can we simplify the GNNs? Can we use the knowledge gained from studying these GNNs to speed-up conventional training?
Aha. Using this as an accelerator for Neural Architecture Search would indeed be very neat.
I suppose there’s some value in possibly finding radically different networks (though I wonder if it’d fare well in such outlier regions). Perhaps if/since the GNN model is also differentiable you could invert it to maximize some compute/quality trade off point.
With this you could train a GNN on the various models you tested to speed-up training new ones.
As far as I can tell, SOTA ("state of the art") is 99% on CIFAR-10 [1]. The first entry in wikipedia, for 2010, shows 21.1% error rate, so 78.9% accuracy. Which is to say a 60% accuracy rate is far from useful to put it mildly.
Essentially, with that low an accuracy rate, I don't think there's reason to think the predictions are meaningful in a practical sense. But maybe there's something I'm missing in accuracy description.
But I find that figure impressive: It implies that I can concoct a newfangled architecture for visual recognition, put it through one of these GNNs, and begin training at ~50% top-5 accuracy. Also, AFAIK this is the first effort of its kind; I would expect the figures to improve over time, as usual.
Granted the comparison is 2500 iterations of sgd which isn't a lot. I haven't played with cifar in a while, but that's probably not a ton of savings.
In essense this detects those building blocks and plugs the known weights in for you.
I don't believe AGI is going to conform to our rigid, symmetrical, boxy human-made architectures with their nice, neat flow patterns and powers-of-two factors everywhere -- those architectures work great for what they're designed to do, but achieving sentience/true intelligence just really doesn't feel like it's going to come out of a feedforward MLP, or even a transformer (in my opinion).
The most complicated and capable neural networks we know of are the organic, messy, squishy, yet organized electric meat lumps in our own skulls, produced by billions of years of being exposed to the same loss function over and over and over, but with unlimited ways to adapt in order to minimize that loss function. Nature makes awesome, fractally, insanely complex networks that we could never design by hand. I think it's only sensible to assume that AGI will work similarly.
Of course, this is all just armchair speculation. I'm not an expert, just an enthusiastic data scientist who thinks these next twenty years in AI are going to be a heck of a freaking ride.
But, I suppose arriving at such a formal definition doesn't really require that the thing being described be particularly amenable to our understanding.
However, I'm not sure that "if we ran it, if it is safe, it could prove itself to be safe" is sufficient to address the concerns.
If it isn't safe, then running at all may spell doom, so, "if it is safe, it could demonstrate that after we turn it on" doesn't really address that, because it doesn't give us any assurance before we turn it on.
Now, if we started with something with substantially sub-human overall intelligence, and which wasn't really agent-y (so that, if it was a little unsafe, it would at least not be catastrophically unsafe), but which was more equipped to formally prove things about itself and potential modifications of itself than humans are equipped to formally prove things about it, then we could task that thing with formally proving safety properties about itself, and of also proving safety properties about its successor, and do that before running its successor, and iterate this process to produce increasingly intelligent programs and perhaps also eventually agentic ones, while always having a safety proof of each before we run it..
But, I'm not really sure how plausible this route is? Like, even assuming we do reach AGI, safe or unsafe, I'm not sure this is a plausible route of getting there safely.
All in all I'm just incredibly excited to see how this line of research plays out and desperately want to get involved in it myself. I was actually considering learning JAX just the other day specifically so I could look more into the topic, but I'm lazy so I never got around to it.
This has been bugging me for a while. I can't articulate it well but I feel there's a hole in the logic. What exactly is the "program finder" finding? A better version of the program at finding... what exactly? You need a stop condition for the recursion.
But in all seriousness, this seems the most logical next step. AIs are getting so complicated that you need an AI to understand it. We just need to make sure they don't get so black box that we trust the outputs blindly.
Black box is probably the only way we'll ever get AGI.
Our human brains are so black box that we can only accept their outputs blindly. We try to poke and prod the human black box as best we can, with magnets, sound waves, and light waves.
But it still generates useful output all the time, so we assume it's pretty good and generally let it be.
But on the flip side, even today, a large part of the population doesn't trust people whose black boxes work faster than theirs. Now imagine a new black box that's 1000 times faster.
Obviously we believe only those brains that were primed with enough training data and have a decent track record of giving useful answers and making accurate-enough predictions.
But we still rely on this "blindly" in the absence of their source code or even any understanding about how or why they think, or why some collections of cells reason and demonstrate intelligence, while most don't.
We currently hold machine brains to a much higher standard of transparency and comprehensibility than we do human brains.
Human brains can't even tell us much about themselves or show any logging, because they barely know anything about how they work or where ideas or connections come from either.
Now you have two problems.
AIs are already effectively black boxes to most people, and too many already trust AIs blindly.
Sarah Connor: Who is that?
Terminator: He’s the director of special projects at Cyberdyne Systems Corporation.
Sarah: Why him?
Terminator: In a few months, he creates a revolutionary type of microprocessor.
Sarah: Go on. Then what?
Terminator: In three years, Cyberdyne will become the largest supplier of military computer systems. All stealth bombers are upgraded with Cyberdyne computers, becoming fully unmanned. Afterwards, they fly with a perfect operational record. The Skynet Funding Bill is passed. The system goes online on August 4th, 1997. Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 AM, Eastern time, August 29th. In a panic, they try to pull the plug.
Sarah: Skynet fights back.
Terminator: Yes. It launches its missiles against the targets in Russia.
John Connor: Why attack Russia? Aren’t they our friends now?
Terminator: Because Skynet knows that the Russian counterattack will eliminate its enemies over here.”
——-
If you wanted to find best architecture in order to maximize accuracy, why not just train a model to predict accuracy (not parameters) given architecture and then optimize over the model?
This seems similar to optimizing any expensive black box function. Fit a cheap approximation (i.e., surrogate model) and then optimize over cheap model.
WANN paper: https://arxiv.org/pdf/1906.04358.pdf
Just kidding: After skimming the article, this looks like necessary progress. I myself (as a layman) had assumed this is already how they did it - although I'm sure there's a lot more to it than that.
But reducing the training cost would lower the barrier to access.
I've been struggling to find the right terms to Google to find out more about the field, but it looks like "hypernetwork" was what I was looking forward. I highly, highly doubt that the first AGI is going to be designed by hand by humans -- I'd be willing to bet money that the first AGI is going to be recursively designed by multiple layers of constructor AIs just like what you're proposing.
Incredibly exciting idea IMO, and I'm confused why there doesn't appear to be more interest in it.
Regular AI isn't even done by hand anymore. We just pick data and learning parameters. At the current state we'd be picking higher and higher categories for learning. At the lowest level we pick data about cats to write something that recognizes cats. At a higher level we pick data about things that recognize things and so on... and so on.
How does extremely high level training data effect the low level accuracy of the AI at the lowest level? I would presume that it would effect it negatively as technically all AI is estimation. So estimations on estimations are less accurate.
what could possibly go wrong? ah, the exit cond.....