Why neural networks struggle with the Game of Life (2020)
bdtechtalks.com
bdtechtalks.com
- both show neural networks can learn the game of life just fine
- the finding is that to learn the rules reliably the networks need to be very over-parameterised (e.g. many times larger than the minimal size needed for hand-crafted weights to perfectly solve the problem)
This is not really a new result nor a surprising one, nor does it say anything about the kinds of functions a neural network can represent.
It's an attempt to understand an existing observation: once we have trained a large overparameterized neural network we can often compress it to a smaller one with very little loss. So why can't we learn the smaller one directly?
One of the theories referred to in the article and paper is the lottery hypothesis, which states that a large network is a superposition of many small networks and the larger you are the more likely at least one of those gets a "lucky" set of weights and converges quickly to the right solution. There is already interesting evidence for this.
I feel something similar goes on in us humans. Interesting to think about.
Exponentiation means it is more efficient to start by far exceeding the required reliability and then optimizing the most expensive subsystems/parts. It is less efficient and far more frustrating if multiple things have to be improved to meet requirements.
"The number of synapses in the brain reaches its peak around ages 2-3, with about 15,000 synapses per neuron. As adolescents, the brain undergoes synaptic pruning. In adulthood, the brain stabilizes at around 7,500 synapses per neuron, roughly half the peak in early childhood.
This figure can vary based on individual experiences and learning." -- written by GPT-4o
Confirmed by e.g. https://extension.umaine.edu/publications/4356e/
Or in other words, even in an information theoretic sense, it's true: you can't teach a old dogs new tricks. You need a new dog.
So the program is not just the zeroes and ones, so to speak, but also more nebulous real-time activity, passed on through time. Like a wave on the ocean.
https://news.ycombinator.com/item?id=38505856 https://www.elijahwald.com/origin.html
> The strange thing about all this is that we already have immortality, but in the wrong place. We have it in the germ plasm; we want it in the soma, in the body. We have fallen in love with the body. That’s that thing that looks back at us from the mirror. That’s the repository of that lovely identity that you keep chasing all your life. And as for that potentially immortal germ plasm, where that is one hundred years, one thousand years, ten thousand years hence, hardly interests us.
> I used to think that way, too, but I don’t any longer. You see, every creature alive on the earth today represents an unbroken line of life that stretches back to the first primitive organism to appear on this planet; and that is about three billion years. That really is immortality. For if that line of life had ever broken, how could we be here? All that time, our germ plasm has been living the life of those single-celled creatures, the protozoa, reproducing by simple division, and occasionally going through the process of syngamy -- the fusion of two cells to form one—in the act of sexual reproduction. All that time, ^^that germ plasm has been making bodies and casting them off in the act of dying. If the germ plasm wants to swim in the ocean, it makes itself a fish; if the germ plasm wants to fly in the air, it makes itself a bird. If it wants to go to Harvard, it makes itself a man.^^ #weirding The strangest thing of all is that the germ plasm that we carry around within us has done all those things. There was a time, hundreds of millions of years ago, when it was making fish. Then at a later time it was making amphibia, things like salamanders; and then at a still later time it was making reptiles. Then it made mammals, and now it’s making men. If we only have the restraint and good sense to leave it alone, heaven knows what it will make in ages to come.
> I, too, used to think that we had our immortality in the wrong place, but I don’t think so any longer. I think it’s in the right place. I think that is the only kind of immortality worth having -- and we have it.
Can't say the same about neural networks (yet?).
Isn’t that another way of saying the optimization algorithm used in finding the network‘s weights (gradient descent) can not find the global optimum? I mean this is nothing new, the curse of dimension prevents any numeric optimizer to completely minimize any complicated error function and it’s been known for decades. AFAIK there is no algorithm that can find the global minimum of any function. And this is what currently limits neural network models: They could be much simpler and less resource hungry if we had better optimizers.
I completely agree that the most effective regularization is inductive bias in the architecture. But bang for buck, given all the memory/compute savings it accomplishes, SGD is the exemplar of implicit regularization techniques.
Is it actually proven, or another hypothesis? What is the reason behind this?
In this case a positive benefit of combinatorial complexity.
The lottery hypothesis intuitively makes sense, but as an outsider I find this concept for evaluating learning methods really interesting - To hand craft a tiny optimal networks for simple yet computationally irreducible problems like GoL as a way to benchmark learning algorithms. Or is it more than that? for a sufficiently small network maybe there aren't that many combinations of "correct solutions", so perhaps the way the network emerges internally could really be interrogated by comparison.
Not according to the hype merchants, hucksters, and VCs who think word models are displaying emergence and we're 6 months from AGI, if only we can have more data
Can you do a neural network that, given a starting position of the game of life, decides if it cycles or not? ;)
Ok, not cycles... dies, stabilizes, goes into a loop etc.
The other difference is I don't take it seriously.
https://writings.stephenwolfram.com/2024/05/why-does-biologi...
He says it's possible for smaller games (fewer rules) but unlikely for larger ones.. IMHO anything Turing complete would have this problem.
As it stands, my guess is that the LLM would always confidently make a decision, even if it were wrong, and then politely backtrack if you pushed backed, even if it were originally right.
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Please don't fulminate. Please don't sneer, including at the rest of the community.
I'm not really sure it's the best idea to accuse someone of breaking the rules if in doing so you're also breaking one yourself.
"As the researchers added more layers and parameters to the neural network, the results improved and the training process eventually yielded a solution that reached near-perfect accuracy."
So, no, we aren't asking too much from it. We just need more compute.
They are displaying emergence. They might as well be the walking definition of it.
If the fit was due to a lucky subset of weights you could have train smaller networks many times instead of using many times bigger network.
So it must be something more. Like increased opportunity to create best solution out of large number of random lucky parts.
I think there should be way more research on neural pruning. After all it's what our brains do to reach the correct architecture and weights during our development.
To be clear, I'm not suggesting macro scale, i.e. not quantum, reality itself is probabilistic, only that our ability to interpret perception of it and model it is statistical. That is, an observation or a sensor doesn't actually tell you the state of the world; it is a measurement from which you infer things.
Viewed through this standpoint, maybe the Game of Life and other discrete, fully-knowable toy problem worlds aren't as applicable to the problem of general intelligence as we imagine. A way to put this into practice could be to introduce a level of error in both the hand-tuned and learned networks' ability to accurately measure the input states of the Life tableau (and/or introduce some randomness in the application of the Life rules on the simulation), and see whether the superiority of the hand-tuned network persists or if the learned network is more robust in the face of uncertain inputs or fallible rule-applications.
I would be very surprised if (a) was not effective, but that (b) is difficult is not surprising, since that is a very nontrivial task that requires intermediary modelling tools to perform for humans (arguably the most advanced NN that we have access to at the moment)
(a) is actually a form of (b) in the form of a modelling tool.
Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length
For those who didn't read the article, the content doesn't support the title.
It's a really nice read.
Simple systems built on simple rules creating universally complete computation behaviors are both unintuitive to and underrated by common man.
How is that interesting? It's the definition of the Game of Life... it's not like it's a natural system that you don't know the full rules for...
Surprised that nobody mentioned him yet.
Go’s complexity comes from two players alternately picking one out of a very large number of options.
GoL’s complexity comes from a very large number of nodes “picking” between two states. That’s not precise, just illustrating that there is some symmetry of simplicity/complexity, at least to my eyes.
We could even think of both as collections of 3d structures showing all valid structures possible for a board of size n by n. There are some differences, every single 3d Conway structure has a unique top layer, while Go does not. But that seems like an overall minor difference. There are many more Go shapes than Conway shapes given the same N, but both are already so numerous that I'm not sure that is a difference worth stopping the comparison.