Scrabble is nowhere close to a solved game
medium.com
medium.com
I'm suspicious when I see a result with a one-percent swing like that. How unlikely would that result be if Quackle played 500 games against itself? What would be the average of the score differences at the end of each of those games?
I think the null hypothesis should be that top humans and AIs are already playing perfect games, and we also have to rule out the hypothesis that AIs are consistently better, but by such a small margin that the AIs' abilities are being lost in the random noise of the letters they draw.
Would it be possible to model luck as a separately controlled variable, giving one half of a self-playing AI "helpful" (or unhelpful) letters and seeing how sensitive the results are to that?
The fact that a game with a random element has a coin toss outcome doesn't preclude the possibility that both players are playing close to optimally.
Perhaps a human makes one mistake per hundred games, and a superhuman AI makes one mistake per 1000 games. You could still get a result like 252-248 to the human just from random variance.
To pick an extreme example, let's say that the first player to get an "unlucky" rack gets stuck in loop that somehow keeps them at a disadvantage, one which perhaps snowballs. As an analogy, imagine a game of chess where each player rolls a die at the start of each turn, and if they roll a 6 they have to lose a pawn of their choice. It might be possible for a chess grandmaster to beat a chess computer under that rule set nearly 50% of the time, even with the human making sub-optimal moves in "many situations".
Remember, in a perfectly random game, 50% of the time it doesn't matter if you make a mistake, because you were going to lose anyway. (I don't just mean there are psychological factors of players getting lazy when they know they're unlikely to win, I mean that these mistakes don't affect the outcome, so we're talking about very small statistical effects here).
Anyway, your experience with Scrabble games probably gives you some good intuition for how big an effect luck could possibly have, and I'm happy to accept an assertion that 500 games would be enough to show a statistically measurable lead for a truly perfect player over both the top human and top AI players. It would still be nice to get some data on the effect of randomness, though, for example by measuring the difference in total score between an AI that gets a standard random draw of tiles each turn, and an AI that always draws some set of tiles in the n-th decile of expected score.