Also I think the model is to a large extent "talking itself into" the alignment faking in its scratchpad. These language models are really suggestible, and in my experience prone to self-suggestion as well due to their auto-regressive nature. Section 3.3.3 of the full paper addresses this a bit - there's way less faking when the CoT scratchpad is removed.
Is it still 12% without the scratchpad?
While we’re at it, how can I tell that you aren’t a word salad generator?
Observation of people who suffered from brain injury that resulted in a more or less unmediated flow from the language generation areas all through vocalization shows that we can also produce grammatically coherent speech that lacks deeper rationality
Here I can only read text and base my belief that you are a human - or not - based on what you’ve written. On a very basic level the word salad generator part is your only part I interact with. How can I tell you don’t have any other parts?
My fingers also were involved in the typing of that message, actually they were the last proximal cause of the characters appearing the comment.
Are you saying that on a very basic level my fingers are the my only part you interact with?
But that's not what I think people mean when they say "world salad generator" or "stochastic parrot" or "broca area emulator".
The idea there is that it's indeed possible to create a machinery that is surprisingly efficient at producing natural language that sounds good and flows well, perhaps even following complex grammatical rules, and yet not being at all able to reason
What exactly would be your bar for reconsidering this position?
Taking some well-paid knowledge worker jobs? A founder just said that he decided not to hire a junior engineer anymore since it would take a year before they could contribute to their code base at the same level as the latest version of Devin.
Also, the SOTA on SWE-bench Verified increased from <5% in Jan this year to 55% as of now. [1]
Self-awareness? There are some experiments that suggest Claude Sonnet might be somewhat self-aware.
-----
A rational position would need to identify the ways in which human cognition is fundamentally different from the latest systems. (Yes, we have long-term memory, agency, etc. but those could be and are already built on top of models.)
https://techcrunch.com/2024/12/14/klarnas-ceo-says-it-stoppe...
>> What exactly would be your bar for reconsidering
The salad bar?
A junior engineer takes a year before they can meaningfully contribute to a codebase. Or anything else. Full stop. This has been reality for at least half a century, nice to see founders catching up.
It’s not a matter of opinion that LLMs are autocomplete word salad generators. It’s literally how they are engineered. If we set that knowledge aside, we unmoor ourselves from reality and allow ourselves to get lost in all the word salad. We have to choose to not set that knowledge aside.
That doesn’t mean LLMs won’t take some jobs. Technology has been taking jobs since the steam shovel vs John Henry.
Autocomplete can use a search engine? Write and run code? Create data visualizations? Hold a conversation? Analyze a document? Of course not.
Next token prediction is part, but not all, of how models are engineered. There's also RLHF and tool use and who knows what other components to train the models.
Obviously it can, since it is actually doing those things.
I guess you think the word “autocomplete” is too small for how sophisticated the outputs are? Use whatever term you want, but an LLM is literally completing the input you give it, based on the statistical rules it generated during the training phase. RLHF is just a technique for changing those statistical rules. It can only use tools it is specifically engineered to use.
I’m not denying it is a technology that can do all sorts of useful things. I’m saying it is a technology that works only a certain way, the way we built it.
“We’re Axel & Vig, the founders of Innate (https://innate.bot). We build general-purpose home robots that you can teach new tasks to simply by demonstrating them.
Our system combines a robotic platform (we call the first one Maurice) with an AI agent that understands the environment, plans actions, and executes them using skills you've taught it or programmed within our SDK.
If you’ve been building AI agents powered by LLMs before, and in particular Claude Computer use, this is how we intend the experience of building on it to be, but acting on the real world!
…
The first time we put GPT-4 in a body - after a couple tweaks - we were surprised at how well it worked. The robot started moving around, figuring out when to use a tiny gripper, and we had only written 40 lines of python on a tiny RC car with an arm. We decided to combine that with recent advancements in robot imitation learning such as ALOHA to make the arm quickly teachable to do any task.”
The problem is that this knowledge is an a priori assumption. If we're exercising skepticism, it's important to be equally skeptical of the baseless idea that our notion of mind does not arise from a markov chain under certain conditions. You will be shocked to know that your entire physical body can be modelled as a markov chain, as all physical things can.
If we treat our a prioris so preciously that we ignore flagrant, observable evidence just to preserve them -- by empirical means we've already unmoored ourselves from reality and exist wholly in the autocomplete hallucinations of our preconceptions. Hume rolls over in his grave.
My favorite part of hackernews is when a bunch of tech people start pretending to know how very complex systems work despite never having studied them.
It's such a basic part of the field, I'm doubtful if you're talking about me in the first place? Nobody who has ever interacted with the intersection of quantum physics and computation would even blink.
[1] - Refer to Mathematical Foundations of Quantum Mechanics by Von Neumann for more information.
[2] - Refer to Quantum Chromodynamics on the Lattice for a description of a QCD lattice simulation being implemented as a markov chain.
Really? Is that what’s wrong with me? ¯\_(ツ)_/¯
People get very hung up on this "autocomplete" idea, but language is a linear stream. How else are you going to generate text except for one token at a time, building on what you have produced already?
That's what humans do after all (at least with speech/language; it might be a bit less linear if you're writing code, but I think it's broadly true).
It might actually be linear — how minds actually function is in many cases demonstrably different to how it feels like to the mind doing the functioning — but it doesn't feel like it is linear.
Nobody even knows where, specifically, qualia exist in order to be able to direct technological advancement in that area.
Is this you trying to humblebrag that you are such a talented writer you never need to edit things your write?
But ideas are not. The serialization-format is not the in-memory model.
Humans regularly pause (or insert delaying filler) while converting nonlinear ideas into linear sounds, sentences, etc. That process is arguably the main limiting factor in how fast we communicate, since there's evidence that all spoken languages have a similar bit-throughput, and almost everyone can listen to speech at a faster rate than they can generate it. (And written text is an extension of the verbal process.)
Also, even comparatively simple ideas can be expressed (and understood) with completely different linear encodings: "The dog ate my homework", "My homework was eaten by the dog", and even "Eaten, my homework was, the dog, I blame."