It’s a next letter guesser. Put in a different set of letters to start, and it’ll guess the next letters differently.
It’s a next letter guesser. Put in a different set of letters to start, and it’ll guess the next letters differently.
https://www.anthropic.com/research/tracing-thoughts-language...
> Instead, we found that Claude plans ahead. Before starting the second line, it began "thinking" of potential on-topic words that would rhyme with "grab it". Then, with these plans in mind, it writes a line to end with the planned word.
My guess is that they have Claude generate a set of candidate outputs and the Claude chooses the "best" candidate and returns that. I agree this improves the usefulness of the output but I don't think this is a fundamentally different thing from "guessing the next token".
UPDATE: I read the paper and I was being overly generous. It's still just guessing the next token as it always has. This "multi-hop reasoning" is really just another way of talking about the relationships between tokens.
Interpreting the relationship between words as "multi-hop reasoning" is more about changing the words we use to talk about things and less about fundamental changes in the way LLMs work. It's still doing the same thing it did two years ago (although much faster and better). It's guessing the next token.
At least in my view it's still inherently a next-token predictor, just with really good conditional probability understandings.
It shows that we, computer scientists, think of ourselves as experts on anything. Even though biological machines are well outside our expertise.
We should stop repeating things we don't understand.
I feel that we pick the next thought to convey. I don't feel like we actively think about the words we're going to use to get there.
Though we are capable of doing that when we stop to slowly explain an idea.
I feel that llms are the thought to text without the free-flowing thought.
As in, an llm won't just start talking, it doesn't have that always on conscious element.
But this is all philosophical, me trying to explain my own existence.
I've always marveled at how the brain picks the next word without me actively thinking about each word.
It just appears.
For example, there are times when a word I never use and couldn't even give you the explicit definition of pops into my head and it is the right word for that sentence, but I have no active understanding of that word. It's exactly as if my brain knows that the thought I'm trying to convey requires this word from some probability analysis.
It's why I feel we learn so much from reading.
We are learning the words that we will later re-utter and how they relate to each other.
I also agree with most who feel there's still something missing for llms, like the character from wizard of Oz that is talking while saying if he only had a brain...
There is some of that going on with llms.
But it feels like a major piece of what makes our minds work.
Or, at least what makes communication from mind-to-mind work.
It's like computers can now share thoughts with humans though still lacking some form of thought themselves.
But the set of puzzle pieces missing from full-blown human intelligence seems to be a lot smaller today.
We simply haven't.
Are we just now rediscovering hundred year-old philosophy in CS?
All that means is that treating something as a black box doesn't tell you anything about what's inside the box.
I ... did you respond to the wrong comment?
Or do you actually think the DB table can genuinely reason about things?
Good luck.
For a very vacuous sense of "plan ahead", sure.
By that logic, a basic Markov-chain with beam search plans ahead too.