> The OP is claiming that they have created an accurate world model for Othello, so I think discussion of intentional mistakes and thermodynamics is outside the scope.
It's not out of scope because these all fall under the problem of induction, which is what I mentioned in my first post. There is no such thing as achieving "certainty" in such scenarios where you don't have direct access to the underlying model, there is only quantified uncertainty. This was all formalized under Solomonoff Induction.
So I'm making three points:
1. Your requirement for "knowledge" to be "100% certainty" in such scenarios is just the wrong way to look at it, because such certainty isn't possible in these scenarios, even in principle, even for humans that are capable of world models, eg. swap a human in for GPT and they'll also achieve a non-zero error rate, even if they try to build a model of the rules. There is no well-defined, quantifiable threshold at which "quantified uncertainty" can become "certainty" and thus what you define as "knowledge". Therefore "knowledge" cannot be equated with "certainty" in these domains. The kind of "knowledge" you want is only possible in cases when you have direct access to the model to begin with, like being told the rules of Othello.
2. Even if you do happen upon an accurate model, you'd never know it and so have to continuously retest it by trying to violate it, so you cannot infer the lack of a internal model from the existence of an error rate. The point I was making with this is that your argument that "a non-zero error rate entails a lack of an accurate world model" is invalid, not that GPT necessarily has an accurate world model in this case.
3. I also dispute the distinction you're trying to make between statistical "pattern matching" and "understanding". "Understanding" must have some mechanistic basis which will have something like this kind of pattern matching. I assume you agree that a formalization of Othello's rules in classical logic would qualify as a world model. Bayesian probability theory where all the probabilities are pinned to 0 or 1 reduces to classical logic. Therefore, an inferred statistical model that asymptotically approaches this classical logic model, which is all we can do in these black box scenarios, is arguably operating based on an inferred world model with some inevitable degree of uncertainty as to specifically which world it's inhabiting.