Ironically, I think he was right! In fact some of his initial experiments, like copycat, were about predicting patterns. You can see next-token-prediction from there. But I think he always held out for an algorithmic/logical method rather than a purely statistical one.
If he had accepted the "Bitter Lesson", I think he would have been at the forefront of LLMs.