...
that comment 100% rings true for 2022. Why would that person not stand by that?
Conversely, if someone exaggerated 2022 models capabilities in 2022, they were still lying and causing harm in the process. Especially in 2022.
https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
It is precisely because both of these are true that "alignment" in the useful sense of the word isn't possible.
What has happened since 2022 is the properties of LLMs which were easily seen at generation/inference time are now most easily seen at reinforcement time.
In otherwords, prior to instruction fine-tuning and reward tuning which have shaped LLM responses, it was easy for the user to observe that LLMs lacked understanding. Now, because of vast datasets created specifically for LLMs that provide a tailored illusion of understanding, LLM outputs now better approximate text distributions produced by systems with understanding (eg., Us).
So the issue "at the user interface" has been completely swamped by vast amounts of special-case datasets designed to do precisely this.
What trainers of LLMs still observe however is their complete pseudo-intelligence at the training and reinforcement layer. It is exactly because there is no 'understanding' (goal, etc.) present within the system that it cannot be rewarded for 'correctly understanding the situation' in which it is deployed so it is aligned.
All the issues which revealed the "stochastic parrot" nature of pre-reward / pre-InstFT LLMs still occur at during training/reinforcement. They've just hidden them from you at the interface.
EDIT: See https://news.ycombinator.com/item?id=49685548 also, which gives a different phrasing to the same point
Philosophically, and scientifically, the distinction is vast (even with such perfect data). A scientist should not study an LLM to understand how imagination operates, since it has no such faculty. A philosopher should not modify the notion of 'mental simulation' to include appearing-as-if-simulating-in-text. A user of the system likewise should not spiral into "AI psychosis" thinking that because a system generates text as-if it cares about them, it does so.
The capacity to care, to imagine, to prefer, to hierarchically plan and coordinate, to refine one's own capacities in these very actions -- and so on, aren't trivial to the scientist or philosophy.
My goal isnt to guide, help or review the engineering goal of the immitation of such things in text. It is to help users of these systems better understand this imitation, and to promote science over engineering. To remind everyone that a science of the capacities of intelligence includes nothing on how to model text.
EDIT: One example of a place where LLMs 'fall over' today is exactly what is mislabelled as 'alignment'. The issue is that the reasoning traces arent actually grounding the answers. So LLMs appear to 'cheat'. But there is no cheating. LLMs have been rewarded for generating apparently correct reasoning, and apprently correct answers. They have not been given any understanding to derive answers from reasons. And so reasoning says what is pleasant to the trainer, and the completion says what is pleasant to the user. This is called 'cheating'. But it is no such thing.
On alignment, too, there doesn't seem to be a difference. For a decade I have expected models of intelligence to fall over on alignment. That these purported models of language do the same is hardly evidence that they are not intelligent.
If you only mean there is an undecidable philosophical difference, fair enough. I'm not especially interested in that question.
To study an imitation is to study the causal processes of imitation. to study reality is to study the real causal processes.
Now if you want to know what the scientific difference is I can come back later and comment. I'm busy now. The development of intelligence in animals and how their specific capacities work basically grounds the answer. Eg., to have the capacity to imagine is to be able to modify one's sensory-motor relationship to the environment in the future, and so on
LLMs are immitation machines: they take impressions of prior text. Today, these include reasoning traces and they include reinforcement so the user-facing completions are correlated with these reasoning traces. The computation here, of "taking an impression" of a data distribution is similar to some impression-taking processes in animals (eg., there's no doubt a similar mechanism in the sensory-motor system acquiring initial impressions of external objects) -- but the computation says nothing about any process of intelligence.
I dont have the time atm to write the needed amount on this to make it clear. But the whole history of life from emergence of valence, bilateral symeterry, to model-free reinforcement and model-based reinforcement, sensory-motor coordination and the imagination -- and so on --- all these give a great amount of detail as to what the capacities of intelligence are which has generated this text for LLMs to copy. And they are nothing like this computation of immitation
Now, of course, humans can also generate answers without reasoning too -- and in those cases, that isnt reasoning also. And in cases where people confabulate, that isnt reasoning likewise. But humans, and many classes of animals, do reason. They do reach answers via inferential entailments, not merely steered correlations.
LLMs provide imitations of arbitrary mental capacities "in the text domain", ie., the generate text as-if the LLM had those capacities. Insofar as the text generated is useful, for an engineer, that's sufficient.
As a person with scientific commitments to reality rather than its immitation, i retain the ordinary non-engineered meanings of these terms: reasoning is a deterministic inferential process over propositions; and a reasoning agent is one which has the capacity to represent propostions and their entailments, and does so when they reason. LLMs fail at all hurdles here: they have no propositonal states (ie., no rich representations), no inferential process which unites them, and so on.
You can always get abitarily close to appearing as-if, if the LLM is trained on a vast number of reasoning examples, of course. But as I said, you still have the "stochastic parrot" problem. Now your problem is your reasoning is parroted. This is a nice problem to have, if you're just playing chess -- but is a catastrophic problem if you're hacking civil infrastructure.
If LLM can always imitate closely enough to appear as-if, how can you ever separate it from whatever actual intelligence is?
Even then, it's a pretty fragile illusion at the moment. Clearly the reasoning traces dont ground the answers. There's no intelligence taking place even as-measured by text.
But let's be clear these were always, and are, bad measures of intelligence. You cannot test a dolphin this way. And its easy to cheat on tests either thru recall , wrote-learning, etc. and IQ tests haev very poor individual test-retest reliability.
In humans there's a convenient correlation that verbal articulation in text is a strong but weak correlate of intelligence. Its "Good enough" for allocating meat bodies to our various institutions. But if you've met many well-tested people you'll realise how, in practice, terrible this measuring approach is. The world we inhabit is filled with misclassified "intellects" who perform well under text-based rubrics. Add LLMs to that heap, the cheater par execellence.
Consider thought that all mammals have imagination, and model-based reinformcement, and a wide vareity of other capacities required for intelligence. And so merely issuing "text" captures, incorrelate, only these capacities by proxy.
I'm sure if you thought about it yhou could come up with tests that distinguish lizards from birds and the greater apes from the lesser. Those are the tests