Which incidentally would also imply that a lot of N-grams just don't exist in the training data, causing the model to completely halt when someone says something unexpected.
But that's not the case - instead we convert words into a lot fuzzier float vector space - and then we train the network to predict the next fuzzy float vector. Since the space is so vast, to do this it must learn the ability to generalize, that is, to extrapolate or interpolate predictions even in situations where no examples exists. For this purpose, it has quite a few layers of interconnects with billions of weights where it sums and multiplies numbers from the initial vectors, and during training it tries to tweak those numbers in the general direction of making the error of its last predicted word vector smaller.
And since the N-gram length is so long, the data so large, and the number of internal weights is so big, it has the ability to generalize (extrapolate) very complex things.
So this "probability of next word" thing has some misleading implications WRT what the limits of these models are.