When we teach children how to do arithmetic, we have them predict missing items in equations. We don't accuse them of "only doing autocomplete". The same applies for LLMs.
When we teach children how to do arithmetic, we have them predict missing items in equations. We don't accuse them of "only doing autocomplete". The same applies for LLMs.
I'm not making a "popular mistake", I'm literally describing how inference is done.
You are very clearly not doing that. Nothing about your comment had anything to do with the internal structure of LLMs.
I believe you that you've set up some models with pytorch or whatever, but this seemingly hasn't translated to a sufficiently coherent mental model to make the distinction between the extrinsic optimization criterion and intrinsic behavior.
We feed in a context and it gives a probability distribution for the next word. We sample from that distribution following some set of rules. We then update the context with the new word and repeat.
That algorithm is an autocomplete algorithm. As long as that's what LLM inference looks like, all problems that we want to feed to an LLM therefore must be translated to autocomplete.
I think you're making the mistake of assuming that because I use language that resembles the language used by cynics I'm therefore arguing that LLMs are useless. I'm not. All I'm saying is that we need to have an accurate mental model for the way these things work, and that mental model is autocomplete. Nearly every major failure in an LLM application was the result of failing to keep that in mind.
I don't disagree with this, but I do disagree with this earlier statement:
> The further you get from autocomplete, the less reliable the resulting product
Any naturally sequential problem is trivial to translate to autocomplete with minimal loss of fidelity.
Again, I think you're putting words in my mouth and thoughts in my head that aren't there. A lot of people have reacted to AI hype by going the other way and underestimating them—that's not me. I think there are lots of problems they can solve, I just think they all boil down to autocomplete and if you can't boil it down to autocomplete you're not ready to implement it yet with an LLM.
This is completely wrong.
It's just completely wrong to say everything in LLM land is autocomplete. It's trivial and common to do the above.
That's a very clever way of reducing a problem to autocomplete, but it doesn't change the paradigm.
It's a silly way to think about it. Have you seen how people are fine tuning for classification? It's not like fine tuning for instruction or summarization etc, which are still using next token prediction and where the last layer is still mapping to vocab_size outputs.
To make it even more concrete.... I have an LLM where I removed the last layer and fine tuned on a classification problem and the last layer now only has two outputs rather than vocab size outputs. The goal is binary classification. The output is not a completion of an idea or anything of the sort. It's still an LLM. It still works the same way all the way up until the last layer. The weights are the same all the way up to the last layer. It works because an LLM has to create a rich understanding of how a bunch of concepts work together.
Using the same analogy: I dont see this in LLMs, I can speak with an LLM months and if we dont release a new version it is the same.
So they are not like children and I am not sure we really understand learning in children more so to the level that we can compare LLM with a child way of learning.
We know how LLMs work fundamentally and what their limits are. LLMs are only able to make “correct sounding” statements which have the side effect of being correct a certain percentage of the time. They do not have the ability to reason nor engage in high level thought.
That is technically true but deceptive: that subset is enormous! Skim any HN thread and you'll see many people talking about reasoning ability or how we just need a little more magic sauce and the "hallucination problem" will be solved. And a lot of what people say in this vein is not even wrong.
And that's just HN. Non-technical users are of course going to assume that if it looks like a duck and is dressed up as a duck by its creators, then it's a duck. Why wouldn't they? So I would claim that the subset is a majority.
We know how they work only at the lowest level (the arithmetic operations) and the highest level (the optimization criterion and the representation of various layers, like the input/output layer and for things we can easily probe like embedding matrices).
We do not know "what they are doing" on the inner layers. This is an area of active research.
> They do not have the ability to reason nor engage in high level thought.
You are speculating (and probably incorrectly), or you're really holding back some valuable research from the field of AI interpretability.
You along with others here are far overstating the unknowns we have within the context of AI, whether this is the result of a misinformation campaign targeted at trying to boost the value of this tech or if the pop-sci takes have really gotten too prevalent, it is unclear to me.