Do I understand correctly that it's brute forcing a small grid rather than learning the algorithm?
If by small grid you are referring to the attention matrix plot shown, then that is not a correct interpretation. That diagonal-like pattern it learns, is 3x3 convolution, so it can compare the neighbours of a given cell.
Edit: and note that every grid it is trained on / runs inference on is randomly generated and completely unique, so it cannot just memorise examples
Using 2D RoPE instead would in principle allow scaling up as well, and maybe even period detection if you train it across a range of grids, but would eventually hit the same issues that plague long-context scaling in LLMs.
Would it be possible to train an LLM on the rules how we would teach them to a human?