> This is like saying that we know code is correct (bug free) because most of the lines sort of look correct. It's not /inherently/ an untrue statement, but it doesn't match up with any sort of real world experience.
That's actually a really good analogy. I'm thinking so long as the code "builds", then it's likely that the code can execute, and at that point, the code is worth reviewing.
There's only around 600 characters/words to guess across about 700 tablets. Let's say there's about 2 sentences per tablet, so roughly around 1400 sentences to test against the guessed dictionary.
If we simply loop through, iteratively building up this dictionary, making small adjustments and happen to get to something like 80% of the sentences passing the review (i.e. not nonsense sentences, saying something with meaning, thew meaning makes sense in the context of the period etc). Then our word predictions are likely close to the original.
Like what are the chances you can choose the wrong word across lets say 200 sentences, and still maintain meaning in all of those sentences? And then multiply that guessing all the words?