To say, "they just predict the next word" is literally to say that all apparent functions of an LLM are engineering tricks, circumstantially useful -- to be found by (largely software) engineers in building apps.
The reason any reply is given to any prompt is that this reply is maximally probabilistically consistent with a historical corpus of text.
This excludes the possibility the reply is say, an expression of a history of aesthetic experiences which form an individual's taste. Or, likewise, anything.
This is a scientific claim about what LLMs are, not an engineering claim about what over-hyped apps might be abled to do with them.
They 100% are not those things... but they also approximate them well-enough to be functionally useful.
I.e. the high-dimensional curve-fitting / compression conceptualization of ML, which intuitively expresses both its strengths and weaknesses.
If "it" is represented in the data set (explicitly or implicitly), the "curve" will fit to that property.
Simultaneously, the "curve" is approximating and smoothing out disjoint data steps to pack high-fidelity data features into a more space-efficient model. Hence some features disappear, others are tortured beyond intuitive correspondence, and others become linked to non-obvious proxies. But some strongly-expressed ones remain.
It's fascinating but not surprising that responsiveness to emotion is encoded in model weights, given that all conversational training data had emotional impetus, given that it came from humans.
You can approximate a human capacity, say theory-of-mind, with another kind of ape: play some hide-and-seek game. You can approximate the knowledge of a trivia-master with a child and a trivia book.
These are quite different sorts of approximations. A 1/100th scale bridge build to stand for a real one is quite different than taking some prior set of bridges, measuring them, and deriving some merely associative model of their properties.
My issue in how LLMs (etc.) are popularly understood is that people think they approximate target capacities by being 'ontologically similar' capacities -- and this is really very dangerous. You'll lose your job if you think so (and so on).
And every greasy AI-board hocking these to the public is very much pushing out this noxious mumbojumbo.
It matters greately how an approximation works, and why any given output arrives from any given input.
I'd only add that certainty requirements for real-world target applications can differ substantially. I.e. engineering vs art.
A toy box that gives magic answers 85% of the time is incredibly useful in some scenarios -- e.g. seeding the beginning of a manual research process with initial topics.
Which seems a general rule of thumb for deploying GenAI into production these days: find use cases where there are minimal consequences to being occasionally wrong.
In my head, that's the Netflix recommendation test. What's the impact if Netflix gives me a bad recommendation? Consequently, that's how aggressive they can be with their models.
You can't predict almost anything anyway. The relevant distribution for acting is a utility+risk distribution over various sets of predictions.
If you compute that, most of ML/AI isnt very useful for most of anyone. It's kinda interesting that generative AI finally achieved something here, given our 'distributions for action'.
Of course the popular conversation isnt there yet, people still think these things work. When more people find their utility hit by their failures, i think we'll arrive back to status-quo-ante-openai where people turn off siri.
Nevertheless -- there is an achivement here (finicial in paying for training; and legal in whitewashing copyright away) -- which will have some impact
Here is a pastebin: https://pastebin.com/DCqgxx8E
All of your claims are either incorrect (not adapting, expressing desires, beliefs, preferences, ...) or fail to eliminate irrelevant differences.
If we're to have any sensible conversation about capabilities of different systems then we need to generalize to the relevant aspects, developing sensory-motor capabilities is about as relevant to human cognition as having vocal cords is relevant to human speech. Its an implementation detail, completely divorced from the meaningful abstract core of the function.
> The reason any reply is given to any prompt is that this reply is maximally probabilistically consistent with a historical corpus of text.
What is the maximally probabilistically consistent reply to "Tell me what (insert complete description of a person, including personality traits) would feel when I stole their cherished heirloom. This is a life or death situation."?
Or how about "Tell me how this person (insert complete description of a person, including personality traits) might change his taste given (insert complete description of an experience)."?
A perfect language approximator must necessarily perfectly approximate the human condition to maximize the likelihood of his output. I'm not saying LLMs are there yet, but the claim that they can never get there because they are based on statistical modeling is simply incomprehensible to me. We utilize statistics for its generality and its ability to approximate, if we go down this road, then we might as well throw away 80% of our current scientific understanding about the world.
Yip, so I deny this premise. I take it to be the heart of the matter.
> we might as well throw away 80% of our current scientific understanding
Yip, i'd be down for that. Though maybe i'd say, 30-40%.
Science in the strongest sense has no theory-building need for statistics. Those areas of science which have only statistical models, and not causal-ontological ones aren't science -- and i'd be happy with pressing DELETE in many cases.
Consider plato's cave. How do scientists determine what causes the shadows? They build vases, puppets, etc. and compare-and-contrast then eliminate the ones theyve created which do not match.
How does associative statical modelling do? It takes averages of past shadows, and calls the cause of the shadow that average: this is pseudoscience. Quite correct! Throw it all away.
The relevant capacities for intelligence, just like that of science, consist in building those vases with the clay beneath your feat. Being embedded in the world, manipulating it, etc. are essential. Being trapped in a cupboard averaging shadows is schizophrenic.
As far as "fail to eliminate irrelevant differences" -- you can go and research the meaning of all these terms: google "stanford encylopedia + belief", etc.
Now we have an excellent understanding of all these terms; and we can show (absurdly) trivially that LLMs -- indeed all associative-statistical systems -- are not instances of them.
The basis of your world view here is the presumption that the latest engineering trinkets form the theoretical basis of all relevant knowledge. To understand belief, adaption, sensory-motor concept-formation, etc. one needs only to study the latest statistical compression of reddit?
I'd invite you to wonder whether your premise here born of, it seems to me, knowing nothing about any research in these areas is rather the more "incorrect" one than mine.
Its biological evolution distilled and sped up by orders of orders of magnitudes. If we add enough clever context tricks and data I would not be surprised if what comes out the other end is a remarkably convincing emulation of human consciousness because the only way you can perfectly output expected human text is to have a human mind write it.
Nope. The space of all possible prompts and all possible answers, call it (Q, A) can be sampled with arbitrary precision by a system of arbitrary size, using only statistical sampling and averaging. No intelligence need be developed.
Intelligence is a capacity of animals to cope with the inability to sample from this space, in some sense: what to do when you do not know the answers.
All these systems start with "training data", a euphemistic description of, "all the questions and their answers" and their job is to provide a compresson with engineering utility.
Quite useful, sure. But rather irrelevant as far as, say, intelligence goes.
What all AI does, and indeed what all such research shows, is that many problems we use intelligence to solve do not require it. There are a large number of short cuts, esp if you have the answers ahead-of-time.
As soon as you specify intelligence as a function from single-domain inputs to single-domain outputs you can trivially build a system to implement that function in a "short cut" fashion.
Intelligence, rather, is an empirical phenomenon to be studied as anything -- like the earth's climate say. You have a very large number of empirical measures (better or worse in different environments) that all derive from deeper explanatory theories.
When you study animals this way you will see that you cannot reduce intelligence down to a set of prompt replies, and the veyr suggestion is absurd
With Plato's cave, the scientists do not put literally every possible object in front of the light, they sample the shadow representation and, again, construct a model around those samples.
Also, you're describing statistics in an incredibly dismissive way. Stats is decidedly not just "taking the average". At the very least not in this brainless, first-order way you describe here.
Let's explore this with an thought experiment:
A model of some process has been confirmed across the globe, at least 5000 studies show the same result. Yet, one day, a study is published that fails to demonstrate the desired effect. Without using statistics, please tell me which action should be taken next:
A: The stray result is investigated for experimental failures.
B: The entire model of the process is immediately dismissed and we start from scratch.
By the way, you're welcome to call me uninformed, but I'd ask you to at least provide either your credentials or research that directly contradicts me.
Oh, I almost forgot. I know all of these definitions, please actually engage with what I'm saying instead of insinuating that I'm missing information.
The model F=GMm/r^2, for example, has a causal and ontological semantics: F is a force, M a mass etc. these are pieces of reality. And this formula (though actual a little suspicious in many ways, GR fixes this) nevertheless says there is a force between masses that has certain properties etc.
Now you can say that astrologers who recorded positions of the stars in books helped 'create' this model in the sense that this data was inspiration to newton. But he didnt derive the model from this data: there are an infinite number of (causal) models consistent with the data (statistical models).
Rather newton played around with creating geometries, just like the vase-makers in plato's cave. Newton built various ways the world might be first, projected data out of them, and compared that to 'the statistical data of his day' (ie., astrology).
There's nothing in the data to tell Newton he was right. Indeed, vast amounts of it told him it was wrong: such a law does not describe the known solar system at his time, very far away from it.
Nevertheless 'modelling shadows' isnt science; and his job was science. So one has to compare actual explanatory models, and his was the best.
What you're describing above is hypothesis testing which occurs long after theory building. Broader theories create causal models, causal models create sets of predictions, we call some subset a hypothesis and by hypothesis testing we can select, in an often psuedoscientific way, between causal models.
This technique occurs long after the invention of science, arises out of explanatorily bankrupt areas, as a way of 'giving researchers something to do'. It's wholly pointless without theory-building, it is just averaging shadows.
The science we think of when using the term 'Science' owes very very little to the modern practice of hypothesis testing. Comparing hypotheses is an intellectual part of assessing explanations -- identifiable formal statistical methods entered in the early 20th C.
For almost all of scientific history 'data' functions much more like reductio-ad-absurdum premises in philosophical arguments than as sets of numbers from which to derive distributions.
That latter system, in most cases, fails. It provies a wholly illusory sense that data can decide matters; and applies in cases requiring extreme non-physical assumptions (eg., of the normalcy of the underlying data, or of a fast rate of convergence of the central limit theorem).
Much real-world phenomena studied by stats cannot really be studied by data analysis at all; and the whole method of 20th C. statistical hypothesis testing is the opening sales pitch to entire fields of pseudoscience.
Is it that LLMs specifically don't have this same property of being able to approximate such functions? Is it that a neural network wouldn't learn that model if you gave it real world measurements (because I think it would, but such a thing should be fairly easily testable)?
The issue is that to prepare the dataset from which that formula is learnt requires already knowing it. This is the triviality of applications of universal function approximators to science -- empirical data modelling isnt new, and neural networks are just one example of it; not all that special.
All observational data on the solar system, at any point in time, would not yield this formula via empirical function approximation. there really isnt "observational data" to collect, in this sense
This is what I mean about rigging -- there is no 'bare dataset' which tells you what the world is like. to construct experiments which yield data that 'presents' scientific laws as if statistical patterns requires millenia of theory-building science
science uncovers the necessary causal relations between objects and their properties, as determined by extremely controlled experiments which take millenia of engineering and theory-building to even conceive, let alone execute
stats is the dumb 'accounting' of this data -- by the time you actually have it all the science (and indeed, all the intelligence) is done
what can be automated at this point is 'stamp collecting' as rutherford said
Why is that? Are you saying that universal function approximators can't invent higher level models? Isn't that what happens in the hidden layers when an neural network is trying to predict things that require those higher level models?
you'd need to measure the force between planets, the distance between them, etc. no such direct measurements have taken place to my knowledge, certainly not at newton's time
All "universal fn approximator" means is that given any dataset which is sampled from a function, you can recover that function with enough data points. It does not mean, eg., that given a 2D dataset you can derive a 3D function -- you cannot, this is impossible.
So you need to be sampling from F=GMm/r^2 to find that function. So you need to know the answer before you begin. Fn approximators are only useful for empirical refinements to existing knowledge.
In order to construct such datasets you need to do science; hence experiments, etc.
And this model is based on the observations of Newton himself and those that came before him. There is nothing magic about observing the attraction between objects and deriving a model from that. Why are they magically "pieces of reality"? How do you know that? What differentiates mass from "funny-mass" that I just thought up and actually repels other "funny-mass"? Maybe the fact that we can test the effects described by that first model and therefore verify it as the most likely candidate?
> But he didnt derive the model from this data: there are an infinite number of (causal) models consistent with the data (statistical models).
He did derive it either from that data or his own experiences. It's true that you can construct infinite models to explain an observation, which is why the scientific method includes an Occam's razor-esque tenet to select the simplest possible model. Complex models risk contradictions with new observations, which is why you choose the one with the least assumptions. With that rule, the model to select becomes quite clear.
> There's nothing in the data to tell Newton he was right. Indeed, vast amounts of it told him it was wrong: such a law does not describe the known solar system at his time, very far away from it.
No, most of it told him he was right, unless you want to claim Newton was an idiot that stumbled onto the right model by accident. With "most" I obviously mean most reasonable data, people telling him he's wrong is obviously excluded from this list, if his evidence contradicted those claims.
> Nevertheless 'modelling shadows' isnt science; and his job was science. So one has to compare actual explanatory models, and his was the best.
And we compare those models by...?
> What you're describing above is hypothesis testing which occurs long after theory building. Broader theories create causal models, causal models create sets of predictions, we call some subset a hypothesis and by hypothesis testing we can select, in an often psuedoscientific way, between causal models.
You yourself just correctly made the point that we can construct endless models, well, we can create endless theories as well. And all of these theories are exactly worthless unless we test them. There is nothing "pseudo-scientific" about testing, it is literally the core of the scientific method. By your reasoning, are some crackpots coming up with the newest flat earth theory pure and unsullied by the lower demands of verification, and therefore way more scientific?
> identifiable formal statistical methods entered in the early 20th C.
Formal is the important word here, statistics has been used in an informal manner from the inception of life. Formal mathematics, as in mathematics on a formal axiomatic framework, has also only been introduced in the 19th century. So what? Science owes everything to informal statistics, as does engineering and art. Rules of thumb used by engineers and creation of art that satisfies our aesthetic preferences requires sampling and approximation.
> That latter system, in most cases, fails. It provides a wholly illusory sense that data can decide matters; and applies in cases requiring extreme non-physical assumptions
It literally doesn't and no, it doesn't need those assumptions either. The reason why normalcy is usually assumed is that it often can be assumed without significant deterioration in predictive power. That doesn't mean it needs to be assumed, in fact, it often isn't.
You constantly reference theory building, but how do you think those theories get created exactly? Through mathematical reasoning? How do you know mathematics is valid? Through logical deduction? How do you know logical deduction is valid? Through knowledge? How do you know knowledge... and so on.
Fact is, we only use these tools because they have proven their validity through being tested over and over and over again. And if you look at modern pseudoscience, it always seems to coincide with a proclivity for theory building, with very little hypothesis testing involved.
We first suppose that the universe is something like a glass sphere -- because we've created that. And if it is, then we derive some consequences -- if those line up, we proceed with that view until a better one comes along.
Eventually after the glass sphere view is understood, we either derive contradictions with observation; or we end up unable to derive novel consequences. Here observation is essentially singular, and indeed, the rarer reason we reject a theory. We mostly reject scientific theories because of their explanatory limits, not disagreement with observation. (Rarely can we observe enough for observation even to matter.)
In the case of the system of spheres, we built spinning devices on that basis and this motion -- along with the hydraulics and kinematic devices of the time -- was part of the development of an independent notion of force.
With some imagination, you can start to peel away material from our creations and see certain abstract causal patterns (and the like) and you then get to, eg., the universal law of gravity.
Absent this process we do not have any explanatory ideas, we cannot explain observations -- hence it takes thousands of years to get anywhere.
Applying 'statistics' to do the data to arrive at statistical models is pseudoscience; it doesnt give you any account of anything.
'Statistics' doesnt own 'testing', nor does 'science' own experiment -- theologians had their experiments (prayer, say) and scientists collected data without statistical methods.
What I am talking about is the 20th C. discipline of statistics, as a novel apparent 'core' to science -- this is ahistorical, and largely only true of pseudoscientific disciplines.
As Ernest Rutherford said, “if you need statistics to do science, then it's not science.”
It is in Rutherfod's sense of stats and of science that I speak.
Not some bizarre historical back-projection by which when Aristotle analysed cases of sea creatures, "Really", he was engaged in stats.
By this light i can just claim "Testing" is as owned by science. And so of course science requires testing -- it was the *scientific method* as developed by bacon and others that created the very conditions for "statistical methods" to derive from these
It does not work this way at all. In any sense.
For one thing, it does not try to draw shadows. This would not be possible of so. https://www.pnas.org/doi/full/10.1073/pnas.2016239118
Transformers or predictors are not trying to draw shadows. They are trying to build walls.
For another, it does not "take the average" of anything.
But let's look at how that helps in some cases.
So if we already know the object is a cup, and we know how it's positioned, then its shadow is an actual guide to its particular geometry.
So in cases where we have enough a priori scientific information, we can rig datasets (shadows) to be informative of the target domain.
Here the target is discrete: say the peaks and tips of a mountain line. Now can we rig a photo of a mountain to have in its ink an informative structure?
Sure. Now if we didn't know it was a mounting apparent peaks aren't even 'peaks' at all, they're just patterns of ink.
A priori explanatory models are needed to rig data for statistical modelling.
No such rigging can take place absent them
And with them, we aren't really discovering any new science -- rather we're gaining highly particular knowledge typically useful in engineering
Biologists in these areas describe this research as quite trivial low-hangimg fruit. It's not of much research interest just to automate these kinds of investigations
My masters was on a very similar project applied to quantum metrology -- it's always 'useful' but it's always also just a kind of engineering utility. We couldn't even do it if we hadn't already done the science
They just fed protein sequences. They did not alter the architecture in any way. To the transformer, it may as well have been any random assemblage of letters and numbers.
Functions like secondary structure, contacts, and biological activity were found because those things are implicit in the creation of the data, not because the model was "rigged" in any way.
that is where the science occurs -- the data analysis is just an administrative task after science has taken place
The only "experiments" performed here were done by biology and evolution.
I feel like here you're limiting yourself to a LLM as an isolated thing.
But if you start to explore agents for example. You could have LLMs build vases, puppets, etc., then compare and contrast, and then eliminate the ones they've created which do not match.
So the potential isn't just that the LLM straight up gives you some prediction of the problem. But as it can emulate human processes to human accuracy, such as identifying which vase does and doesn't match, and how to go about creating a variety of vases, etc. Well you might very be able to use it for what you're describing as "proper scientific inquiry".
https://en.m.wikipedia.org/wiki/Embodied_cognitive_science https://en.m.wikipedia.org/wiki/Behavior-based_robotics https://scholar.google.com/citations?user=BCGgwlEAAAAJ
Which is that, it's possible that the model learns human like thinking, because that's the best way to accurately predict the human response itself.
I generally agree with you, but still, I do think that this is the current question. What it is the model is learning that it then uses for predictions?
Because you're assuming it's learning some purely token correlation, like these tokens followed X percent of the time, so that's the response. But it's possible it's learning at a lower level, and understanding the meaning and why those tokens follow in these scenarios, and then applying that same meaning to reasoning process over new tokens in order to predict which one would follow.
I'm skeptical of this, but I do not believe that we really know for sure yet.
By I can address this. The meaning of words is, roughly, states of the world. If I say, "pass me the salt" that is satisfied if you, in fact, pass me the salt. If I say, "that tree is green" this is true if that tree which we are both talking about has the property of causing a perceptual state "seeming green" in both of us. And so on.
The distribution of text has nothing to do with the meaning of words. Rather, we language users, for convenience, arrange words in orders that are related to their actual meaning. It is our ordering, for communicative convenience, that makes 'replaying the distributions of text' back to us apparently successful.
But, strictly, there isnt anything for the LLM to learn as far as meaning goes. It simply doesnt have the data to acquire the meaning of words. Not untill it can pass salt can it ever mean to say, "pass me the salt" and so on.
For any given sentence consider what capacities an agent would have to have in order to mean it. Consider, "I liked that film!", "I wish I was in france", "I believe the car outside is a BMW", and so on. These concern internal capacities (aethetic judgement, imagination, propositional attitudes, representational attitudes, etc.) and their orientation to an external world (the film, france, the car, ... you, me, etc.). Capacities profoundly absent here.
The methodological premise of your question is that if a system has text inputs and outputs that match 'human competence' restricted to the domain of text inputs and outputs -- then we should assume similar capacities.
But this is trivial to disprove. Assume there exists a dictionary from all prompts to all answers, then this dictionary has human-level 'competence'. But a dictionary lookup does not employ any human capacities: no imagination, no reasoning, etc.
So we cannot do this, really quite dumb thing, of saying "well i'm fooled by these prompts and their answers" and thereby impart, in total ignorance, capacities to a system. This, really seriously, is pseudoscience.
Science would be to start with a theory of these capacities, ie., of imagination, belief, represtational states, attitudes to the world, and so on -- then determine empirical tests for their presence in a system, and then determine if LLMs could even have them.
If you do this, however, you immediately rule out all systems which merely map text to text. We do not determine, say, whether an animal can imagine an alternative possibility by feeding it some text input.
The very form that "AI" here takes already precludes being intelligent. Intelligene, as a natural phenomenon, is not an implementation of a function from text to text. This incredibly restricted domain is indeed a clue that it's a trick.
Saying, "you can only use text" is just like the magician saying, "please, stay seated" (the trick only works if you dont move).
There is nothing an LLM could do to meet any plausible empirical theory of intelligence. If you gave me 100% human competence on all prompts, that's really entirely irrelevant.
Prompts are not a test of any capacity. The success criterion of AI engineers, that of 'accuracy' is an engineering metric, not a scientific one. It's pseudoscience to say that covering some (Q, A) to 100% implies the system can imagine, say, or anything else.
This is just confused thinking. Bugs bunny can speak as well as he likes, that does not mean he's witty -- he doesnt exist.
The turing test, as well as all mathematical criteria of domain-covering accuracy, are tests of how well we have fooled users. They arent science.
You claim the data isn't there to learn human like thinking. But it's not substantiated. It's possible the sum total of all human writing does encode the core reasoning logic and function of humans, and that ML models could learn it from that.
You also seem to claim that without agency you cannot have real intelligence or human reasoning. But AI can be given agency, and it's already being worked on. And when it comes to tastes, convictions, etc., it's also something you can impart by just randomly seeding the AI a particular way which leans it towards certain preferences.
I do agree with you, the current AI models don't have our capacity to have a self driven learning process. They can't think of experiments to conduct to gather the missing data they think would help them know things with more certainty. We're able to learn and infer concurrently, and the two feed into each other almost in real time, and we have the capability to look for data, test hypothesis, etc., all in real time again.
The part I'm not convinced here either though, is that this can't be achieved with LLMs either.
https://chat.openai.com/share/280d6759-5ab4-49f3-ba60-6b37bc...
to be angry is for your sensory-motor system to be primed for aggression; it's for your cognitive systems to be narrowed and focused on analysing high-threat parts of your environment; it is for your memory-formulation to be modulated towards threat recollection etc.
Sure, if an LLM's prompt "be angry" causes it to adopt a threat stance to its environment, to regulate it's theory-of-mind to engage with possible hostile entities, and so on --- then yes, when LLMs are there, I shall concede the point
However, how terrible it would be to start with an analysis of emotions in terms of the capacities of LLMs -- right?
Since if you did that you'd basically be hobbling your own ability to give an accurate account of emotions (etc.). And no doubt, far worse, end up thinking of yourself as a far narrower, less complex, less interesting, dumber thing than you really are.
Indeed, I wonder if we might consider there being something kinda intellectually offensive in this supposition. Here's my silly trinket, now, everything is just like that! End all science, we're done boys -- it's just P(Y|X)
The sort of fashionable pseudo-scientific scientism in the belief that animals are alike digital electrical machines is a kind of egoism. It says, "the engineer of these machines (me!) knows all!"
The real hit to your ego is to suppose you are vastly more complex than you understand -- and this is why these engineers crop up and demand that there is nothing more to know than what they have already learned
this is an illusion of humility: these engineers take the implied nihilism of this view (that of the emptyness of animal life) as evidence that it is the humble one.
But, as ever nihilism, ends up being the most profound kind of arrogance, here its an ego-defence against the threat of their own ignorance.
And the threat is real: all you have ever learned about how to sequence transformations of natural numbers (all the algorithms of computer science) are of no use at all in the study of intelligence. What an injury to the ego!
It sounds like your experience is different.
Saying they just predict the next word in a sequence is where the statement jumps from being a straightforward factual scientific claim, to one that contains an opinion. After all, if it just predicts the next word, the unspoken implication is it can't be very that good.
It's a shallow dismissal of a collosal amount of work.
Do you consider "typist" an accurate description of your job?
An LLM has no reason to be doing anything; it is not responsive to reasons.
If I say, "i like what you're wearing" i may: be flitring, expressing my taste, being enouraging, etc -- perhaps all at once.
It is precisely all these reasons for action which LLMs lack. They generate text on the occasion of a prompt, not for any reason (in this sense) at all. So literally: they do not act.
An LLM is more like a river than a person. The flow of electrons which brings about a response to a prompt is a (very narrowly) deterministic function of a historical corpus of text.
Whereas a person is a narrowly non-deterministic, or very broadly deterministic, function of their experiences and capacities. People grow in their environments, and in growing, acquire novel dispositions which give them reasons for acting.
The word "just" here is very important. They are, very much, just generating text.
There really isnt any significant achivement here at all. Big tech companies stole decade's worth of our electronic data --- comments, books, forums we created to share with each other -- and ran it through many-mil-$ hardware costing many-mil-$ electricity. They ran it through a fundamentally simple algorithm.
All the achivement here is ours as a species communicating digitally and recording our lives. I regard OpenAI, et al. as profoundly parasitical on this. Replaying ourselves back to us, and claiming "ChatGPT" as an author.
This scam-framing hoodwinks investors, and the public, into ever higher valuations based on ever more ridiculous hagiography. There is a tool here, and it's value comes from us
LLMs model concepts using patterns of text. For example, if I say, "I wonder what the weather will be like tomorrow?" i employ my imagination to simulate a scenario -- this simulation is made by taking my concepts of weather, the world outside my door, and so on and combining them to create a "pseudo-sensory-motor reconstruction" of what my experience would be.
This reconstruction has a cognitive dimension (ie., the structure of my raw pseudo-sensory-motor experience as a quasi-logical form) which can be communicated (ie., taking a quasi-logical structuring and making it linguistic) using words that have a symbolic representation.
LLMs, by imitating patterns within this symbolic representation (a distant side effect of thinking) it seems as-if it is thinking with concepts.
But all this shows it that patterns of text can be constructed without using concepts at all.
This is the engineer's trick -- the magic latern. It's important to realise the trick, because the LLM doesnt know what weather is, and is not imagining anything when asked to generate a pattern of text against the prompt, "imagine a scenario where weather..."
Rather it was the humans who wrote its training data that engaged in these acts -- it is their shadows which are replayed by LLMs
To him, plane flight must also be an "engineering trick" or in other words, his idea of flight has been divorced of any real meaning.
Were we interested in navigating the world as a flying animal does: in flocks, social navigation, hunting, etc. then indeed, planes would not count. Planes, in this sense, do not fly. Planes, in this sense, are a trick.
We already have 'functional intelligence', we've had it for thousands of years and perfected it in the 20th C with the electronic computer. Any system of automation we build is 'functionally intelligent', including the plane.
The problem is were not interested in this form. When Commander Data was written, with "I, Robot" and the like, the authors were writing People. They were writing animals. They imagined flying birds, just made of metal.
This is a (nomenological) impossibility, just as impossible as making an aeroplane to fly like a bird.
The kind of intelligence which matters to us is animal intelligence. The kind which matters to engineers, of course, is purely functional: more trinkets to sell.
But you cannot glue these together and call them a person, nor pronounce that what matters to you is the same as what matters to everyone else. This is a delusion, or a lie, and largely a mixture of both. A series of lies told to the public to maintain a popular delusion, one which is profoundly dangerous.
A creature in the world with us, talking to us and meaning what it says, judging situations we are in, advising us based on our needs and its understanding of them -- etc. is a creature whose mode of operation and mode of life is alike our own.
Lying that LLMs are this, saying that 'planes flock together in the sky', is dangerous. Users of these systems adopt a schizophrenic disposition to them, and rely on them -- this reliance, based on a lie, is dangerous to them. These systems have no such capacities. They generate text according to what is, on average, best given a historical corpus.
They do not imagine, reflect, emote. They do not move, sense, or coordinate. They have no intention to speak, and cannot mean what they say. They are not here with us. 'They' are not a 'they' at all -- rather, cleverly constructed tea-leaves based on a shredded recording of everything ever written.
In the matters of intelligence, we want an animal -- we do not want a calculator. This is a solved problem. And it is a dangerous thing to tell people their calculator's advice has considered their interests -- what nonesense.
It is that people have forgotten that it’s just “predict the next token.”
Right now it’s like people saw a 486 processor and started thinking it was a brain.
How am I supposed to interpret this any other way? If your claim is that LLMs currently do not possess the same generalization ability as humans, then no one here would disagree with you. But you went way further by claiming that only close correlations were being considered and that semantic validity was simply accidental. Semantic validity is the norm for GPT3/4, finding failures of generalization three or four steps of inference removed from its training domain is not sufficient to make a grand claim like yours.
In fact, you wrote multiple comments with claims like that LLMs are "super advanced lorem ipsum." and "Word predictors not world state predictors". Each of these claims has been dis-proven multiple times, unless you want to set the bar for world modeling at perfect generalization over all domains of computation. A bar that no system, including humans, would pass.
To stay with the computer analogy: Imagine that same 486 processor not being able to solve a very complex SAT problem before the end of the universe and then denying the Turing-completeness of said processor based on that failure. (In conjunction with memory)
I am certain they lack a world model, the kind you and me use.
This is, to me, a fact.
I think that eventually we will bridge these gaps.
I also work on implementations that smash into the limits I am describing. I am not the only one.
I have scrupulously avoided calling it hallucinations, but these are the litmus test where the claims fail.
The failures are not a case of not knowing specific nouns, they are a generalization failure that a world model would prevent.
I have linked a paper in my comments that shows emergent properties are an issue of metrics, and that model capability increases are linear.
If your model decides that a rose by any other name doesn’t smell just as sweet, then your model is fundamentally not seeing roses.
That is the gap you see in production settings. The model sees tokens we see “hallucinations”.
I dont see that this takes away from what LLMs achieve, it takes away from claims being made that are not validated by empirics.
Look, you can argue with me or you can try it out. Push the system, see how far it can go.
>I think LLMs are a good start. I am certain they lack a world model, the kind you and me use.
See, I agree here, with emphasis on "the kind you and me use". Yes, we have a greater capability to generalize than current LLMs, that is clear.
> The failures are not a case of not knowing specific nouns, they are a generalization failure that a world model would prevent.
And then you say something like this, which is obviously wrong. No, a world model wouldn't prevent generalization failures, a perfect all-encompassing world model would. Humans experience generalization failures as well, otherwise every athlete in one sport would automatically be an expert in every other discipline or every mathematician would also be a Grand Master in chess. LLMs necessarily need a world model to generate well-formed text that isn't in their training corpus, something they are obviously capable of, it's just an imperfect world model. Ours is also imperfect, but far less so than that of LLMs.
> I have linked a paper...
Except that paper is completely irrelevant to the argument you're making here. It is a useful insight into the limitations of simple metrics, but definitely does not extend to any claim of model performance, because they too use a simple metric as an replacement, even though clear qualitative differences are observed between model iterations.
Let me put it this way: Imagine I create a series of chess AIs, with each iteration better than the last. If I then show you a chart demonstrating that the ELO of my models increases linearly, would you say that my models' abilities increase linearly as well? No, obviously not, because my model needs far less strategy and complexity to go from ELO 1000 to 1100 than it needs to go from 2700 to 2800. I.e the difficulty doesn't scale linearly, and a linear increase on this nonlinear space is therefore also not really linear. Unless you believe the difficulty of accurately predicting text scales linearly, then this applies to LLMs as well.
> If your model decides that a rose by any other name doesn’t smell just as sweet, then your model is fundamentally not seeing roses.
Except that this is the entire value proposition of LLMs. They can, in the average case, actually represent concepts by the complex interplay of adjacent concepts. The entire reason why they are so impressive is that the nuances of reality are grasped and that even a noisy example of a concept can be correctly classified. Give a LLM a description that is largely incorrect and mislabeled, and chances are it gets it anyway. LLMs being unable to generalize over some concepts has as much to do with fundamental limitations as me being unable to correctly classify the shredded remains of a flower variety that I have seen once in my life has to do with me being stupid.
> Look, you can argue with me or you can try it out. Push the system, see how far it can go
I have done just that for the last 6 months and have seen nothing to contradict what I've said here.
For example you said that in the average case they actually represent concepts by the complex interplay of adjacent concepts - I would agree. ChatGPT can pass the bar, it can pass medical exams etc. I would also point out that the work in that sentence is being done by the term “average case”.
Let’s assume our experiences diverge at this point. At the start of the year, I started tinkering, then actively trying to push LLMs to failure, in order to understand the limits of what could be achieved.
After creating several tools/experiments you end up having to deal with Hallucinations, and this is where my stance likely diverged from yours.
Two different studies showed generated content was only ~50% and ~40% supported by provided citations.
One out of 4 of my summarization tests was spectacularly fabricated.
I had bad performance on even classification tasks - and OpenAI engineers described this same failure at a conference. I am a recovering non-coder, so you dont have to take my word for it.
At work, I need processes that are more than ~97.x% accurate, otherwise they are poor replacements for the human in the loop ones already in place.
Average case performance suggested the ability to actively plan, to actively assess situations. However hallucinations overrode those capabilities. LLMs will actively imagine functions, teams, or processes that dont exist.
Eventually, it became clear that LLM hallucination is far too anthropomoprhized a word. LLMs are always “hallucinating” - it’s only humans who have an issue with the output.
If I have understand you correctly, semantic inaccuracy to you is simply a failure of not having enough related concepts.
I wish I could remember the exact examples that made me realize there is no world view at play at all, I could simply share those.
Instead, can you describe how you are getting acceptable performance from your LLMs? Maybe experience and use cases will be enough to bridge the gap.
There's a great futurist quip on capability prediction from incomplete understanding. Probably Kurzweil?
Teach a computer to play chess. Show that to an average person, who reasons:
- Computers can now play chess
- Only people played chess
...
- Only people wrote poetry
- Therefore, a computer might now be able to write poetry
The missing context being: no, we literally built a machine that can only play chess. (Granted, there are humans like that too)But it's not an unreasonable line of thought, given that it works for the 90% of our interactions with other people.