LLMs are pretrained to maximize the probability of predicting the n+1 token given n tokens. To do this reliably, the model learns statistical patterns in the source data and transformer models are very good at doing that when large enough and given enough data. It is therefore suspect to any statistical biases in the training data because despite many advances in guiding LLMs, e.g. RLHF, LLMs are not sentient and most approaches to get around that such as the current reasoning models are hacks over a fundamental problem with the approach.
It also doesn't help that when sampling the tokens, the default temperature of most LLM UIs is 1.0, with the argument that it is better for creativity. If you have access to the API and want a specific answer more reliably, I recommend setting temperature = 0.0, in which case the model will always select the token with the highest probability and tends to be more correct.
That is what they did with these "chain of thought" models. Maybe they didn't do it in the optimal way, but they did train them on their ability to answer certain questions.
So the low hanging fruits from this style of training has already been plucked.
Grab the best model you’ve got access to (o4, Gemini Pro 2.5, Claude Sonnet 3.7, etc) and try playing chess with it. The results are astonishingly bad.
It’s not just that LLMs make poorly thought out moves. They regularly make completely illegal moves (making a knight move like a pawn, for example), hallucinate new pieces into existence, change pieces from one color to another, etc.
More words have probably been written about Chess than any other game. You can literally buy a book that’s just about chess openings. There has to be a staggering amount of text about chess in the training data set. And yet, even the best models are far worse at chess than I was in 3rd grade (and I’m not great).
They do pretty well for the first few moves, often playing well known chess openings. But once they get to the part where you need to reason, they fall apart in the most outrageous ways. I suspect this reveals something about how they deal with other logic-based tasks. I’ve wondered for a while now if they basically code through using a vast amount of training data from Stack Overflow, which is why they’re so good at helping with common error codes, and so useless when doing something novel.
These new chess experiments have been eye-opening about how incapable the best models are of basic reasoning. Unless something about that changes drastically, I think LLMs have pretty much plateaued. There will undoubtedly be some advancements, but LLMs are never going to reach AGI without reasoning, nor will they be able to do most jobs.
For example, LLMs cannot test their thoughts against external evidence or other knowledge they may have (such as logic) to think before they output something. That's because they are a frozen computation graph with some random noise on top. Even chain of thought prompting or RL-based "reasoning" are just a pale imitation of the behavior we actually wish we could get. It is just a method of using the same model to generate some context that improves the odds of a good final result. But the model itself does not actually consider the thoughts it is <thinking>. These "thoughts" (and the response that follows them) can and do exhibit the same defects as hallucinations, because they are just more of the same.
Of course, the field has made some strides in reducing hallucinations. And it's not a total mystery why some outputs make sense and others don't. For example, just like with any other statistical model, the likelihood of error increases as the input becomes more dissimilar to the training data. But also, similarity to specific training data can be a problem because of overfitting. In those cases, the model is likely to output the common pattern rather than the pattern that would make sense for the given input.
Similarly, suppose I always roll dice to determine tomorrow's winning lottery-ticket number. Getting it right one day doesn't change the mechanism I used. Some people might assume I was psychic, but would be wrong.
LLMs is just generation. Whatever pattern they embedded, they will happily extrapolate and add wrong information than just use it as a meta model.
You can do a kind of sensitivity analysis to see how sensitive the output is to small perturbations of the weights... but it's computationally expensive.
Might be an interesting form of fine tuning to do distillation where the student's current sensitivity to noise (extracted from the backwards step of gradient descent) is used to shrink the predicted distribution towards uniform. It could be done very cheaply during training and perhaps could avoid the back propagation cost of doing it at inference time.
https://www.linkedin.com/posts/charlesmartin14_talktochuck-t...
But they are the same thing: the model is extrapolating. It doesn't know when its extrapolations are correct or not because an LLM doesn't have access to the outside world (except via you and whatever tools you give it).
If it was free of extrapolations it would just be a search engine over the training data, and that would be less useful.
It is a single slide, very helpful: https://www.youtube.com/watch?v=ETZfkkv6V7Y
https://www.anthropic.com/research/tracing-thoughts-language...
It observes a so-called "replacement model" as a stand-in because it has a different architecture than the common LLMs, and lends itself to observing some "activation" patterns.
Then it liberally labels patterns observed in the "replacement model" with words borrowed from psychology, neuroscience and cognitive science. It's all very fanciful and clearly directed at being pointed at as evidence of something deeper or more complex than what LLMs plainly do: statistical modelling of languages.
Calling LLMs "next token predictor" is a bit of a cynical take, because that would be like calling a gaming engine a "pixel color processor". It's simplistic, yeah. But the polar opposite of spraying the explanation with convoluted inductive reasoning is just as bereft of substance.