Trivia: Claude Shannon proposed the idea of predicting the next token (letter) using statistics/probabilities in the training data corpus in 1950:
"Prediction and Entropy of Printed English"
https://languagelog.ldc.upenn.edu/myl/Shannon1950.pdf
[1] https://people.math.harvard.edu/~ctm/home/text/others/shanno...
"One may get a remarkable semblance of a language like English by taking a sequence of words, or pairs of words, or triads of words, according to the statistical frequency with which they occur in the language, and the gibberish thus obtained will have a remarkably persuasive similarity to good English."