I've worked with autoencoder, convolutional and LSTM methods with text that appear to solve problems that they shouldn't be able to solve from
https://en.wikipedia.org/wiki/Variety_(cybernetics)#Law_of_r...
that is, they don't contain a correct structural representation of the problem space or they violate the laws of computer science.
For instance, a neural network might do a finite number of calculations to guess at a good next move and play a mean game of chess. This is an O(1) algorithm that can't possibly do things that would take an O(log N), O(N), O(N^2), ... algorithm to do it. (e.g. explore the game tree exhaustively and prove this is a winning move.) Put it in the context of a variable-runtime algorithm such as MCTS which can iterate over the network multiple times and it will crush a grandmaster.
Thus you run into the space of the halting problem, something that pretends to sort a list in O(N) time and gets away with it, ...
People in the field work with poorly defined problems, pick metrics ungrounded in practice, etc. In particular there is no requirement that the results of a neural network be 'correct' they just have to be better than some alternative.
If you have one of these things either working or almost-working you might get curious and try to set up some little experiment based on an intuition at the microscale (e.g. it's obvious how to hand-build the network, tasks the network couldn't possibly learn) and do experiments like this one where you can synthesize unlimited amounts of data...
... frequently you get lost at sea, suspect that the emperor has no clothes, wonder if you got the math wrong somewhere, might find your system performs worse the more training data you throw at it...
... thus these results don't close in a clear story, they don't get published.
The power you describe comes in part from network size (which is why NNs became more useful as we could size up as computing got more powerful/cheaper/more efficient). Small network means less powerful ability. The research showed smaller networks didn't work, but larger ones did. Seems like you'd agree?
I think the whole point was that a small network is capable of solving this problem, but training from scratch without intervention only rarely produced a viable solution.
From the article
- But in most cases the trained neural network did not find the optimal solution, and the performance of the network decreased even further as the number of steps increased. The result of training the neural network was largely affected by the chosen set training examples as well as the initial parameters.
So a solution was clearly possible even in the smaller networks when trained from scratch, it just wasn't likely.
So I agree with your first point - The bigger networks did indeed "figure it out". But I don't agree at all with your second - The smaller networks weren't lacking the power to solve the problem, they were just unlikely to reach the right solution with the available training data.
For example, if we were presented a highly sophisticated repetitive pattern and were tasked to figure the underlying tesselation rules that generate it, we would enumerate rulesets that we know already and then try to find parameters for them with the assumption that those parameters are simple.
Again, my point is that to solve GoL, the ML model needs to be given two things: 1) a few rulesets to choose from; 2) an upper limit on complexity of ruleset parameters.