I read the whole article, but I think this statement largely disproves the hypothesis? We only need a single counter example to show that Othello-GPT does not have a systemic understanding of the rules, only a statistical inference of them.
A model that "knows" the rules will make no errors, a model that makes any errors does not "know" the rules. Simple as.
And I feel this way about much of the article, they state they change intermediate activations and therefore the observed valid results prove that the layers make up a rules engine instead of a statistical engine. And I don't make that leap? Why would that necessarily be the case?
Obviously a statistical engine trained on legal moves will mostly produce legal output. A rules engine will always produce legal output.