A lot of philosophers, mathematicians, scientists, etc. effectively say, "Yeah, everything's just f(inputs of world) -> outputs!" That 'f' is doing a lot of heavy lifting. Which is kind of the point of mechanistic interperability - to make sure we're not jumping ahead of ourselves, and to make sure we're careful when we claim what "deep" and "structure" means, when it pertains to that 'f'.
I’m still hesitant to interpret this as “thinking ahead” without at least seeing some more back-and-forth in the literature first, though. This just seems like one of those spots where it makes sense to give other researchers some time to come up with additional hypotheses to explain the observations instead of focusing on the first one anyone proposes in isolation.
The model would construct/hallucinate pre-conditions to satisfy some final output value that the model was predisposed to.
A lot of that signal could be much simpler stuff. This task is hard. The agent seems stuck. The tests are getting better. The current approach looks promising. All of those things make future success easier to predict without the model actually "knowing what comes next" in any strong sense.
Also, their 25 steps are agent turns, not 25 code edits. The median run had something like 52 steps but only two edits, and the program label stays the same between edits. So "25 steps ahead" may sometimes just mean basically the same codebase, with a bunch of reading and test output in between.
So yeah, I'd say it's consistent with Sutskever's view. But "consistent with" and "confirmatory of" are doing very different amounts of work here.
And that's all it needs. Not reasoning.
Babbage’s Analytical Engine didn’t actually analyze anything, and terminology hadn’t gotten any more clear-cut since.
I suspect exact and/or universal definitions for intelligence, self-awareness, 'feelings' etc will prove to be elusive, and best we'll get is systems/robots etc behaving as if possessing those qualities. With some tests to put a number on them.
Downside is that may apply to us humans too.
One of the biggest things I've learned after the event of LLMs is that humans definitions of intelligence/thinking/reasoning/consciousness/etc are very poorly defined. Not just across society at large, but the sciences themselves.
Something like this is actually a stance in the tradition of inferentialism (see the term sapience). Though "reasoning" isn't like, turing machine computability in this space; from what I understand, it's some abstract notion of the "space of reasons". I don't really understand it, honestly.
There's some merit to this, IMO. When an LLM goes wrong, do you blame the person or the LLM? As in, would you throw said LLM in jail, and hold the LLM accountable? Not right now, at least. I'm not sure if that's what is meant by the "space of reasons", but the intuition is that 'reason' can mean a lot of different things, pragmatically speaking. Reason as a legible audit trail is one of those ways.
But that's arguably getting into the social aspect of 'reason' (important!) and not like, what STEM people traditionally think of as 'reason'.
what would be the difference between not reasoning and reasoning not like a human?
I ask it to highlight that most people have no clear answer to what we mean by reasoning. It's a notoriously hard term to define with an precision, and even harder to define by precision in a way that would exclude LLM's, which is usually what the people I ask it of want to do.
I don't think your steelman version really "works" for that purpose because the immediate response would be that it makes "reasoning" largely irrelevant if you define the term that restrictively.
Mistaking chatbot lookahead as reasoning comes with being gulled by the "artificial intelligence" sales pitch.
They are narrow and limited, but still primitive forms of reasoning.