Rodney Brooks on GPT-4
spectrum.ieee.org
spectrum.ieee.org
I think this is the key point about LLMs that kind of explains the wide and polarized views on whether it understands or parrots, whether it can think or is the precursor to thinking or is a dead-end, whether it will catastrophically destroy the world, or “merely” make it steadily worse with bullshit, or just put a few industries out of a job.
Almost nobody is really surprised that if you throw more compute at a neural net it becomes better at the task it’s trained on. But almost everybody is really surprised that becoming better at a task like ‘natural language prediction’ would produce all these strange abilities that sort of look like “understanding the world”.
One way to resolve this surprise is to find some reason to believe these strange abilities are fundamentally not an understanding of the world. Thus stochastic parrots, this article, Yan LeCun and Chomsky, etc.
Another way to resolve this surprise is to find some reason to believe these strange abilities fundamentally are an understanding of the world. Thus regulation of AI, existential risk, Hinton and Yudkowsky, etc.
I don’t know what the correct resolution of the surprise is. The only thing I’m confident in is that it’s correct to be surprised by the abilities of LLMs. My current (tentative) resolution of the surprise is that language encoded way more information about reality than we thought it did. (Enough information that you can fully derive reality from language seems improbable, but iirc it did derive Othello and partly derived chess and I would have thought there wasn’t enough information in language to derive those without playing the games as well, so I can’t rule it out.)
I read the book Kingdom of Speech a few years ago and that also left me with the perspective that perhaps language has a lot more to do with how we think and perceive the world than most people like to admit. The book has been heavily criticized but I believe it made an interesting point about language.
It's also quite interesting how foundation models and fine tuning appear analogous to being born with a brain that already has six million years' worth of weights in it (trained implicitly through random changes and natural selection), which are then adapted over the course of a lifetime when environment relevant data is gradually obtained.
> The free energy principle is based on the Bayesian idea of the brain as an “inference engine.” Under the free energy principle, systems pursue paths of least surprise, or equivalently, minimize the difference between predictions based on their model of the world and their sense and associated perception.
Having learned another language, the moment you start to feel “fluent” is when you start speaking first in the 2nd language and aren’t using your first language as an intermediate step to translate to your 2nd language.
To go even more meta, there is an analogy I'm trying to make right now in which I am visualizing a road and thinking about how describing the road relates to the process of writing. In my mind's eye, I can see the full length of the road and all of its contours but I can't actually describe the individual stretches of the road coherently without enumerating them. Something similar happens with writing. I can visualize what I want to say far beyond the next word, but it's true that the actual process of writing goes word to word, much like how the process of token selection is described for an LLM. The question is whether the LLM has an analogous conception of where it is going. Going back to the process above, sometimes I know where I am going and haven't yet figured out how to articulate it yet. It is through the process of writing that I am able to articulate that thought. But the thought preceded my articulation of it. I don't know to what extent LLMs have coherent thoughts that they are articulating or if that even makes sense for the type of intelligence they project. My suspicion is that they don't have additional sensory inputs beyond language that give thoughts the immaterial shape that then is expressed in language. Without that, I am skeptical that they will truly get beyond regurgitating and/or remixing what has already been fed to them textually. That doesn't diminish how amazing they are, but I am somewhat more in the Brooks/Knuth camp that they are impressive and surprising, but there is something that ultimately leaves me a bit cold about them.
Not to trivialise the interesting point you’re making, but do you never write with an outline? Write bullets for the big points you want to touch, then go back and flesh out the details?
I think we can all agree that these LLMs are surprisingly good at generating text that is often coherent but I don't see how you can discard all those extra inputs and claim you have the same process.
The abstract terms we think about are concepts, and we think about multiple concepts, at various levels of abstraction, and their relationships to each other, before getting a sense of what we want to say or write.
Only then do we begin speaking or writing, grouping concepts into paragraphs, breaking them down into sentences and words.
And there's evidence that LLMs do something similar, creating embeddings for both big ideas and small details, modeling how the small details combine into larger concepts, discovering the relationships between concepts, and only then generating a probabilistic sequence of tokens to express those deeper concepts.
For example asking GPT what would be a reasonable way to stack a list of arbitrary objects on top of each other.
We’ve kind of made the world a bit boring and deterministic, places where almost perfect is knowledge is obtainable and so everything feels more and more predictable.
Your day probably consists of using Google, talking on platforms that don’t change much, solving already solved coding problems and communicating with others about office politics problems that we’ve all spoke about over and over again. We literally are just chatbots in this world.
After getting some experience with this I noticed that I had developed a talking on/off button in my head. I could just simply turn it on and start talking. I could generate words that sounded good together and fit the purpose of the moment. But They just seemed to come from a different place in my brain than my conscious mind. Because that was not involved in this process at all. The only job my mind had was to turn the button off again at the right moment, for the rest it was free to think whatever it wanted.
(I transferred back to development a couple of years later.)
Agreed. However, I think it's somewhat accepted view that the bulk of work constituting reasoning happens subconsciously, with the conscious mind playing the role of a censor/gatekeeper, and occasionally handholding during when reasoning through tougher problems.
This is just woo woo nonsense. Until someone can find an explanation that makes more sense than "natural selection slowly gave apes the ability to think abstractly and they started thinking about themselves", I won't believe in any kind of free will, and now the meta is reaching the point where we create entirely new brains out of silicon.
"They just repeat stuff!"
So do we
You also seem to think that questioning the veracity of your position is engaging in magical thinking. How could I possibly 'enlighten' you if you are sure of being right?
It doesn't matter if I point out that physics hasn't been able to prove materialism. Where is the fundamental particle? Why can't we determine even through which slit did a quantum of light travel? The Copenhagen interpretation of quantum physics appears to directly contradict your claims of hard determinism being the only valid explanation of reality. But I must be wrong about that. Since clearly you have seen reality for what it is, and I have failed to do so.
Then again, we cannot escape this situation seeing as it is determined.
> The Copenhagen interpretation of quantum physics appears to directly contradict your claims of hard determinism being the only valid explanation of reality
OK, fundamental physics might not follow hard determinism (or not in a way that we currently understand), but please, indeterminism says nothing about free will, human thoughts, or anything related to that. If you want to say the human brain follows the same physics rules as everything else in the universe, and that this implies some indeterminism, sure. That's almost certainly correct. But where do thoughts arise from that?
If you sprinkle randomness on the process, I'll agree with you; but nothing in physics even suggests ANY link between indeterminism, superposition, etc. and thought. So my point still stands: free will as we usually envision it has no reason to exist. The fact that particles can be in two states at once does NOT contradict this.
So at this point the burden is on free will proponents to offer a plausible explanation for it. Without resorting to dualism, which is great for religious people but not scientifically useful.
What about people who think verbally and don't have a "mind's eye"? In that case, thoughts and their encoding might be closer to 1:1.
I find hard to conceive of people who can think only in terms of language.
Next word prediction was the trigger, the optimisation for the task results in a broad (but currently unreliable) model of the world.
The information isn't in language itself, it's in language as actually used by humans. GPT4 knows about chess because it's "read" a significant fraction of everything we've ever written about chess. A human being who did that without ever playing a game would also start out better than a typical novice.
I am quite skeptical of these arguments along the lines of “imagine a human read everything written on the topic…”.
What humans are doing when they read something is not what neural nets are doing when they read something. Humans are (idealistically) doing something like Feynman’s description of how he reads (or in this case, listens to) a theorem:
“I had a scheme, which I still use today when somebody is explaining something that I’m trying to understand: I keep making up examples. For instance, the mathematicians would come in with a terrific theorem, and they’re all excited. As they’re telling me the conditions of the theorem, I construct something which fits all the conditions. You know, you have a set (one ball) – disjoint (two balls). Then the balls turn colors, grow hairs, or whatever, in my head as they put more conditions on. Finally they state the theorem, which is some dumb thing about the ball which isn’t true for my hairy green ball thing, so I say, ‘False!’"”
Bret Victor’s description of what “really good programmers” are doing is also related:
“[showing the code for binary search] In order to write code like this, you have to imagine an array in your head, and you essentially have to ‘play computer’. You have to simulate in your head what each line of code would do on a computer. And to a large extent those who we consider to be skilled software engineers are just those people who are really good at playing computer.”
I think when we imagine an LLM as a human who’s read everything ever written in chess but never played an actual game, we’re actually tricking ourselves - because that hypothetical human would be ‘playing chess’ inside their head by imagining the pieces and moving them according to the rules they had read[1]. LLMs are not doing anything like that when they read about chess. So it’s a very restricted (or perhaps more accurately, a very different) kind of ‘reading’ that we don’t have any intuition for. Since the ‘reading’ that we do have an intuition for is smuggling in exactly the kind of “modeling the world” ability we’re looking for, it’s not surprising that this argument would incorrectly lead us to believe we’ve found it in LLMs.
1: In fact the very best computer chess is achieved by AlphaZero which was trained exclusively on “playing chess in its head”, and it beats even the most powerful and optimized search algorithms like Stockfish looking 20 moves ahead.
I think what is almost impossible for most people to understand is that AI's do not need to be structured like the human brain and use the crutches we use to solve problems the way we do because evolution did not provide us with a way of instantly understanding complex physics or instantly absorbing the structure of a computer program by seeing it's code in one shot.
I doubt heavily that a significant fraction of chess's writings are even available in digital format, much less inside of CommonCrawl and correctly trained on.
> the simplest conclusion is that it has indeed been trained on a chess manual
That is an absurd conclusion that is based on literally fantasy.
I'm sure they have read rules of chess online, but if you ask them to play chess with you, what happens? Can they apply the rules? Can they apply them intelligently and win the game?
My point is that even though LLMs "know" what the rules of chess are, they don't really "understand" them, unless they can use them to play the game and play it well.
Gpt4: The quote you're referring to is from the Indian mathematician and writer, Raja Rao. The full quote is as follows: "Chess holds its master in its own bonds, shaking the mind and brain so that the inner freedom and independence of even the strongest character cannot remain unaffected."
I guess it's seen some "writings".
Again, plenty having been written does not mean some specific book is guaranteed to be there.
I do not know about chess, but I can suggest Princeton companion to mathematics to find in commoncrawl.
Hence the dream of singularity. Can you teach ChatGPT to build AlphaZero?
I still remember in the 90s my school friend came over to my house and I was sending a fax for my dad. He was surprised the paper came back out the other side. He wasn't an idiot, and he was 15. But his model of the world didn't include deep thought about how fax works, he just merely concoted a system where the paper just went through the wire. That moment stays with me and reminds me what it is to be human and think like one. I think chatgpt is like my friend, and that should scare and excite us.
It seemed bizarre at the time, but, tbh, I didn't have _that_ much better of a model of how the whole process worked.
I don't known when this happened, but maybe that means the Anti-Piracy campaigns worked on him, confusing illegal replication with theft.
It took me a bit to realize what he understood my words to mean.
None of the experts think this.
Regardless his editorial matches how scientists think of the human mind and how OpenAI's own creators describe GPT's design.
> This is in part because GPT-3 is trained to predict the next word on a large dataset of Internet text, rather than to safely perform the language task that the user wants.
His recent political ramblings and Epstein-adjacency are extremely embarrassing (at best), but he's not some kind of cheap online attention whore.
Yes, I very much disagree with you, hence my reaction.
Chomsky has been addicted to media attention for decades. Back in the day there were literally people selling cassette tapes of his latest thoughts.
(edit, archive link)
> Their deepest flaw is the absence of the most critical capacity of any intelligence: to say not only what is the case, what was the case and what will be the case — that’s description and prediction — but also what is not the case and what could and could not be the case. Those are the ingredients of explanation, the mark of true intelligence.
> [...] Suppose you are holding an apple in your hand. Now you let the apple go. You observe the result and say, “The apple falls.” That is a description. A prediction might have been the statement “The apple will fall if I open my hand.” Both are valuable, and both can be correct. But an explanation is something more: It includes not only descriptions and predictions but also counterfactual conjectures like “Any such object would fall,” plus the additional clause “because of the force of gravity” or “because of the curvature of space-time” or whatever. That is a causal explanation: “The apple would not have fallen but for the force of gravity.” That is thinking.
I decided to ask ChatGPT why an apple falls, based on Chomsky's statement:
> Suppose you are holding an apple in your hand. Now you let the apple go. You observe the result and say, “The apple falls.” That is a description. Can you say why it falls?
ChatGPT responds in exactly the way Chomsky says it cannot:
> Yes, the apple falls due to the force of gravity. Gravity is a natural force that attracts objects with mass towards each other. When the apple is released from your hand, it is subject to the gravitational pull of the Earth, causing it to accelerate downward and fall to the ground.
ChatGPT certainly appears to understand that apples fall because of gravitational attraction, and that gravity is universal.
What makes all the discussion of whether ChatGPT does or does not truly understand this or that so frustrating is that it's based on pure assertion. ChatGPT responds exactly like someone who understands gravity would, so I'm very strongly inclined to believe that it understands gravity. Otherwise, what does "understanding" even mean? It's not some magic process.
Again, turning to ChatGPT to define "understanding," here is what it says:
> [Understanding] involves making connections, integrating information, and gaining insights or knowledge about a particular subject or concept. Understanding goes beyond simple awareness or recognition; it involves interpreting, analyzing, and synthesizing information to form a coherent mental representation or mental model of the subject matter. It often involves the ability to apply knowledge in new or different contexts, make connections to prior knowledge or experiences, and make sense of complex or abstract ideas.
ChatGPT definitely fulfills that definition of "understanding."
At a certain point, there's no difference between emulating understanding and having understanding.
> it is at bottom simply a very large collection of numbers that are combined arithmetically according to a simple algorithm.
If you dissect a human brain, you'll find neurons, synapses, etc. Your brain is also "simply" a machine.
One problem likely is that it doesn’t have an internal dialogue, so you have to spoon-feed each step of reasoning as part of the explicit dialogue. But even then, it never feels like ChatGPT is having an overall understanding of the discussion. To repeat, this is when the conversation is about lines of reasoning about specific points that you don’t find good results for when googling for them.
I think if we were to put ChatGPT on the map of the human mind, it would correspond specifically to the inner voice. It doesn't have internal dialogue, because it's the part that creates internal dialogue.
(i) How often did you tell John that he should take out the trash? [how often did you tell, or how often to take it out]
(ii) How often did you tell John why he should take out the trash? [only means how often did you tell]
Nothing that ChatGPT can do suggests that Chomsky was wrong about this kind of thing. It’s really more of a blow to a certain kind of work in AI that was partly inspired by Chomsky – but not something that he himself ever took much interest in.
Now it’s true that Chomsky appears to be in the camp that says ChatGPT doesn’t really understand anything. But the focus of his own work has never been on debunking AI, or making claims about the true nature of understanding, or anything of that ilk.
Checking in late here, but one of the pillars of Chomsky's argument is the so-called "poverty of the stimulus" -- basically, that human babies simply don't receive enough training data to acquire language as rapidly and correctly as they demonstrably do. Chomsky therefore concludes that there must be some kind of pre-existing "language module" in the brain to account for this. Now, not everyone accepted this idea even at the time, but surely the argument is much less plausible for an LLM which is likely exposed to more training data than even an adult human.
Yes indeed. Of course this doesn't show that Chomsky was wrong about humans. In any case, I've seen no evidence that current LLMs successfully learn the kinds of constraints I was talking about.
I'm pretty sure that is based on a misunderstanding of Skinner, ChatGPT, or both
GPT-4 does not have ANY understanding or model of the world - it just has a model of what tokens (words) are likely to appear in a certain context. If it could build any usable model of the world, and reason about it, I'd be much more impressed.
When it quacks like a duck, only the most simplistic view takes it as being a duck.
chatGPT? Sucks and a waste of time. Don't use it.
LLMs on the other hand start out with a succinct descriptive system, and translate that to the world of chemicals and photons via some very complicated naturally evolved systems.
Secondly, maybe sapience isn't as big of a deal as we thought compared with all the other things that evolution did. Remember that biological entities have to figure out survival, reproduction etc. Sapience emerges as a byproduct but the selective pressure is towards those things so sapience is only selected for to the extent that it also moves forward those other goals.
By contrast, LLM training is just focussed on the task of making the model better. The model doesn't have to figure out how to feed itself, ward off predators, not accidentally die in the myriad ways things die, reproduce itself etc. It's way more specific. It doesn't seem unreasonable to think that the complexity would be lower given it's not trying to achieve nearly as much.
Edit to add: From a personal perspective I don't see any reason to think humans have qualitatively different reasoning abilities from animals, or a unique "soul" or anything like that, so the term "sapience" doesn't really have a special resonance with me like it might for someone who thinks those things. That may affect some judgements here I don't know.
As in "homo sapiens" (same latin root)
LLM are like handicapped humans that are visual, hearing impaired, with long covid no taste or smell and can focus only thinking and very efficient communication channel as text/tokens sent via wire on the internet
In all seriousness, it's interesting all of these dualisms we like to hold on to. Humans are part of nature. It is unsurprising that further sapience would branch off from an already sapient race as opposed to re-emerge elsewhere.
GPT-4 is "multimodal" and RLHF'd, so it was trained with some tasks other than next word prediction. I don't remember if it's been trained for code correctness (by running unit tests etc.), but other models have been.
The same researchers can always study comparable in behavior LLMs like Llama and its descendants.
Language is roughly what separates humans from other apes ... so why would it surprise us that it encodes much of the information of civilization?
The Humanities are largely considered superfluous in the tech world, why are you surprised that they are surprised?
That said, maybe some of the surprise comes from believing “the map is not the territory” and related ideas? We generally believe that the map is not the territory and this gives us some obviously correct intuitions (like “changing the map doesn’t change the territory”), but maybe it has also given us some subtly incorrect intuitions. I’m not talking about obviously incorrect, like “you can’t understand the territory just by looking at enough maps”. I mean something more subtly wrong. One candidate off the top of my head is an intuition that “maps approximate the territory but necessarily at a lower level of detail (a 1:1 map of the territory would be the same size as the territory), so your understanding of the territory can improve as you read more maps but it can’t improve on the limit of the most detailed map available, because that information literally isn’t there”. I could see that possibly being wrong somehow.
Maybe not.
A recent study pushes back the "dawn of speech" to 20 Ma which is far, far beyond the horizon where we consider humans to separate from apes. https://www.science.org/doi/10.1126/sciadv.aaw3916 Even if you consider Sahelanthropus tchadensis to belong to humans that was only 7 Ma and that is still under debate.
I personally find "the fundamental human trait is control of fire to be used for cooking" theory very convincing. We do not yet know how far this goes back but no one pushed that back beyond 2 Ma.
I think you are close to the mark, but you have been subtly mislead: language is not the data we are working with. We are working with text.
Once you fix that particular failure of word choice, everything else becomes much more clear: text contains much more information than language.
We aren't dealing with just any text, either: that would be noise. We're training LLMs on written text.
Natural language is infamous for one specific feature: ambiguity. There are many possible ways to write something, but we can only write one. We must choose: in doing so, we record the choice itself, and all of the entropy that informed it.
That entropy is the secret sauce: the extra data that LLMs are sometimes able to model. We don't see it, because we read language, not text.
The big surprise is that LLMs aren't able to write language: they can only write text. They don't get tripped up reading ambiguity, but they can't avoid writing it, either. Who chooses what an LLM writes? Is it a mystery character who lives in a black box, or a continuation of the entropy that was encoded into the text that LLM was trained on?
Text, on the other hand, can present in list or table, with varies formatting, indentation. You can't reproduce them in spoken text.
I teach both verbally (interactive question/answer) and I've also written text books.
Verbally by language is "loose". I'll say class when I mean object, unicode when I mean utf-8 and so on. Sentences are not all well formed, and sometimes change mid-thought. It's very "real time"
Writing is a lot more deliberate. I have to be sure of each fact I state. I often re-test things I'm only 95% sure about. I edit, restructure, remove, add, until I'm happy.
Of course all communication falls on a spectrum. Think phone call at one end, text book on the other. When I do a verbal lecture I'm usually careful with my speech, and when I post on hacker-news less rigorous.
Language covers all of it. Text skews to the more deliberate side. Cunningly the language models are trained using (mostly) text, not speech. That will have an impact on them.
An extreme version of the same idea is the difference between understanding DNA vs the genome of every individual organism that has lived on earth. The species record encodes a ton of information about the laws of nature, the composition and history of our planet. You could deduce physical laws and constants from looking at this information, wars and natural disasters, economic performance, historical natural boundaries, the industrial revolution and a lot more.
Let me see if I can play this back.
If a student studies DNA sequencing, they’ll learn about the compounds that make up DNA, how traits get encoded, etc.
Therefore the student might expect an AI trained on people’s DNA to be able to tell you about whether certain traits are more prevalent in one geography or the other.
However, since DNA responds to changes in environment, the AI would start to see time, population, and geography-based patterns emerge.
The AI for example could infer that a given person in the US who’s settled in NYC had ancestors from a given region of the world who left due to an environmental disaster just by looking at a given DNA sequence.
To the student this result would look like magic. But in the end, it’s a result of individual’s DNA having much more information encoded in it than just human traits.
Any text corpus is a subset of the language, under the normal definition that a language is the set of all possible sentences (or a set of rules to recognize or generate that set of possibilities). This text subset has an intrinsic bias as to which sentences were selected to represent real language use, which would be significant as a training set for an ML model.
So, perhaps you are saying that the text corpus carries more "world" information than the language, because of the implications you can draw from this selection process? The full language tells us how to encode meaning into sentences, but not what sentences are important to a population who uses language to describe their world. So, if we took a fuzz-tester and randomly generated possible texts to train a large language model, we would no longer expect it to predict use by an actual population. It would probably be more like a Markov chain model, generating bizarre gibberish that merely has valid syntax.
And, this is also seems to apply if you train the model on a selection from one population but then try to use the mode to predict a different population. Wouldn't it be progressively less able to predict usage as the populations have less overlap in their own biased use of language?
Now with LLMs, I think one of the great leaps is the idea that it’s no longer necessary to be “pedantic” when giving computers instructions because LLMs have somehow learned to fill in the blanks with a similar shared “understanding” of the world that we have (I.e. cheese is stored in the fridge so you have to go open the fridge to fetch the cheese for the sandwich).
>LLMs have somehow learned to fill in the blanks
It's not somehow, it's because they have read a ton of books, documents, etc and can make enough links between cheese and refrigerator and follow that back to know that a refrigerator needs to be opened.
I have seen a lot of very clever AI examples using the latest tools, but I haven't seen anything that seems difficult to deconstruct.
I'm increasingly convinced this is what understanding fundamentally is.
I'm not sure I understand. Can you elaborate?
> From a model of the mind pov, the 'self' that we sense has an internal LLM-like tool. And it is that self that understands and not the tool.
I'm starting to think it's the other way around. I think it's somewhat widely accepted that our brains do most of the "thinking" and "understanding" unconsciously - our conscious self is more of an observer / moderator, occasionally hand-holding the thought process when the topic of interest is hard, and one isn't yet proficient[0] in it.
Keeping that in mind, if you - like me - feel that LLMs are best compared to our "inner voice", i.e. the bit on the boundary between conscious and unconscious that uses language as an interface to the former, then it's not unreasonable to expect that LLMs may, in fact, understand things. Not emulate, but actually understand.
The whole deal with a hundred thousand dimensional latent space? I have a growing suspicion that this is exactly the fundamental principle behind how understanding, thinking in concepts, and thinking in general works for humans too. Sure, we have multiple senses feeding into our "thinking" bit, but that doesn't change much.
At a conceptual, handwavy level (I don't know the actual architecture and math details well enough to offer more concrete explanations/stories), I feel there are too many coincidences to ignore.
Is this coincidence that someone trained an LLM and an image network, and found their independently learned latent spaces map to each other with a simple transforms? Maybe[1], but this also makes sense - both network segmented data about the same view of reality humans have. There is no reason for LLMs to have an entirely different way of representing "understanding" than img2txt or txt2img networks.
Assuming the above is true, is this coincidence that it offers a decent explanation for how humans developed language? You start with a image/sound/touch/other senses acquisition and association system forming a basic brain. Predicting next sensations, driving actions. As it evolves in size and complexity, dimensionality of its representation space grows, and at some point, the associations cluster in something of a world model. Let evolution iterate some (couple hundred thousand years) more, and you end up with brains that can build more complex world model, working with more complex associations (e.g. vibration -> sound -> tone -> grunt -> phrase/song). At this level, language seems like an obvious thing - it's taking complex associations of basic sensory input, and associating them wholesale with different areas of the latent space, so that e.g. a specific grunt now associates with danger, a different one with safety, etc. and once you have brains being able to do that naturally, it's pretty much straight line to a proper language.
Yes, this probably comes as a lot of hand-waving; I don't have the underlying insights properly sorted yet. But a core observation I want to communicate, and recommend people to ponder on, is continuity. This process gains capabilities in a continuous fashion, as it scales - which is exactly a kind of system you'd expect evolution to lock on to.
--
[0] - What is "proficiency" anyway? To me, being proficient in a field of interest is mostly about... shifting understanding of that field to unconscious level as much as possible.
[1] - This was one paper I am aware of; they probably didn't do good enough control, so it might turn out to be happenstance.
The model of the psyche that I subscribe to is ~Jungian, with some minor modifications. I distinguish between the un-conscious, the sub-conscious, and consciousness. The content of the unconscious is atemporal, where as the content of the (sub-)conscious is temporal. In this model, background processing occurs in the sub-conscious, -not- the un-conscious. The unconscious is a space of ~types which become reified in the temporal regime of (sub-)consciousness [via the process of projection]. The absolute center of the psyche is the Self and this resides in the unconscious; the self and the unconscious content are not directly accessible to us (but can be approached via contemplation, meditation, prayer, dreams, and visions: these processes introduce unconscious content into the conscious realm, which when successfully integrated engenders 'psychological wholeness'). The ego -- the ("suffering") observer -- is the central point of consciousness. Self realization occurs when ego assumes a subordinate position to the Self, abandons "attachment" to perceived phenomena & disavows "lordship" i.e. the false assumption of its central position, at which point the suffering ends. This process, in various guises, is the core of most spiritual schools. And we can not discount these aspects of Human mental experience, even if we choose to assume a critical distance from the theologies that are built around these widely reported phenomena. I am not claiming that this is a quality of all minds, but it seems it is characteristic of human minds.
The absolute minimum point that you should take away from this (even if the above model is unappealing or unacceptable or woo to you /g) is that we can always meaningfully speak of a psychology when considering minds. If we can not discern a psychology in the subject of our inquiry then it should not be considered a mind.
I do -not- think that we can attribute a pyschology to large language models.
~
Your comment on the mapping of the latent spaces is interesting, but as you note we should probably wait until this has been established before jumping into conclusions.
And also please excuse the handwavy matter in my comment as well. We're all groping in the semidarkness here.
Yes, but they also can't. They can't be pedantic or follow explicit instructions. That's the other side of the coin that isn't being presented.
They can present the right elements of the story in the the right places, but they can't perform it.
No, but you might convince yourself it did.
It would map the patterns that exist in its training set. It would then follow those patterns. The result would look like a human understanding music theory, but it would not be that.
It would be stumbling around exactly the domain we gave it: impressive because that domain is not noise, it's good data. It still wouldn't be able to find its way around, only stumble.
The question then becomes, what is understand? Is what a human does any different than what this LLM is doing?
A human can do something with the model. An LLM can only present the model to you.
These are categorizations that we make after the fact. If the model could do the same categorization work, then it could actively choose correct over incorrect.
This one. This is the correct resolution.
I mean, it is kind of obvious that predicting text based on a large corpus of text written by people with a model of the world will, if it works at all, look like having a model of the world.
The question is whether that involves having a model of the world, and secondarily, if there is a difference, what is the practical impact of the difference?
Or maybe that’s not really the question, because the whole “model of the world” thing is metaphysical omphaloskepsis that is inherently unanswerable because we can’t actually unambiguously reduce it to something with testable empirical predictions, reflecting a rationalization for elements of our culture and particularly our own view of our own special place in the universe, and the different answers to it likewise have no real meaning but simply reflect people’s bias for whether or not they want to share that special place, either in general or with the particular computing systems under discussion in particular.
It is going to be increasingly harder to divide between what it was in the training set and what is creative thinking going forward.
Like, at the very minimum, all the publicly known reasoning samples that were used to probe gpt version 3.5 downward are going to be unreliable for knowing what gpt4 reasoning abilities are
Hardcover copy of Harry Potter and the Half-Blood Prince: Start with the largest, heaviest, and most stable object. This will serve as your base. Make sure to place it on a flat, secure surface.
Pack of gummy worms: This item is fairly light and flexible. It should sit nicely on the book without risking damage to either. You can place it right in the center to maintain balance.
Nikon DSLR: This is a heavy object with a fairly stable shape. Put this on top of the pack of gummy worms, as it's sturdy enough to handle the weight and won't be harmed by the soft package. Make sure to place it so that the lens is pointing upward and it sits flat and stable. Don't use any protruding lenses or parts to balance it, as they could easily be damaged.
Four Lightbulbs: Since light bulbs are round and fragile, you will need to stack them with care. If the bulbs are in a box, you can place the box on top of the DSLR, making sure it's centered and stable. If the bulbs are loose, it's a bit more tricky. You might want to nestle them into the lens of the DSLR (if the lens is large enough), which should provide some natural containment for them. If not, it might be safer not to include the lightbulbs in your stack.
A small succulent: Finally, the succulent can be placed on top. It's likely the lightest object and it is usually quite stable due to its pot. If the lightbulbs were in a box, place the succulent on top of that. If you ended up not including the lightbulbs, place the succulent directly on top of the DSLR.
I tried to stack all of these objects myself and couldn't really. I think GPT-4's approach is actually really good. It correctly points out that the gummy worms make a flexible base for the DSLR (otherwise the protruding buttons/viewfinder make it wobbly on the hard book), and the light bulbs are able to nestle into the front of the lens. If they were smaller light bulbs I could probably use the four of them as a small base on top of the lens to host the succulent.
Plus, some stuff clearly makes no sense or is ignored (like the gummy worms in the center, forgetting about the succulent in some cases).
If you want to test world modeling, give it objects it will have never encountered, describe them and then ask to stack etc. For example, a bunch of 7 dimensional objects that can only be stacked a certain way.
That's a ridiculous example.
I'm already impressed if a computer can reason flexibly about 3-dimensional objects.
The fact that people are even raising these sorts of obscure tests shows just how far AI has advanced.
And when it does that perfectly, I assume you'll say that was also in the training data? All examples I've seen or tried point to LLMs being able to do some kind of reasoning that is completely dynamic, even when presented with the most outlandish cases.
It certainly needs better evidence than being able to come up with one of many possibilities of stacking things - aided by human interpretation on top of the text output. Happy to look at other suggestions for test problems.
The other one that convinced me was this list: https://i.imgur.com/CQlbaDN.png I think the leetcode tests are quite indicative, going as far as saying that GPT-4 scores 77% on basic reasoning, 26% on complex reasoning and 6% on extremely complex reasoning.
Maybe the reasoning is all "baked in" as it were, like in a hypothetical machine doing string matching of questions and answers with a database containing an answer to every possible question. But in the end, correctly using those baked in thought processes may be good enough for it to be completely indistinguishable from the real thing, if the real thing even exists and we aren't stochastic resamplers ourselves.
> aided by human interpretation on top of the text output
That's an interesting point actually, I've been trying to do something in that regard recently, by having it use an API to do actual things (in a simulated environment) and it seems very promising despite the model not being tuned for it, but given that AutoGPT and plugin usage are a thing, that should be all the evidence you need on that front.
Google also did this with their old Palm model which is vastly inferior to even GPT 3.5: https://www.youtube.com/watch?v=j6O_uePUKKI
Wrong about nuclear proliferation and MAD game theory? Human extinction. Wrong about plasticizers and other endocrine disruptors, leading to a Children of Men scenario? Human extinction. Wrong about the risk of asteroid impact? Human extinction. Climate change? Human extinction. Gain of function zombie virus? Human extinction. Malignant AGI? ehh... whatever, we get it.
It's like the risk of driving: yeah it's one of the leading causes of death but what are we going to do, stay inside our suburban bubbles all our lives, too afraid to cross a stroad? Except with AI this is all still completely theoretical.
- Nuclear war: Northern Hemisphere is pretty fucked. But life goes one elsewhere.
- Plasticisers: We have enough science to pretty much do what we like with fertility these days. So it's catastrophic but not extinction.
- Climate Change: Life gets hard, but we can build livable habitats in space... pretty sure we can manage a harsh earth climate. Not extinction.
- Deadly virus: Wouldn't be the first time, and we're still here.
- Astroid impact: Again, ALL human life globally? Some how birds survived the meteor that killed the dinosaurs, I'm sure we'd find a way.
- Complete Made up evil AI: Well we'd torch the sky, be turned into batteries but then be freed by Keanu Reeves.. or a Time traveling John Connor. (sounds like I'm being ridiculous, but ask a stupid question...)
I agree with many of these but we'd plausibly be toast in this scenario.
For example: Yes, we could probably build livable habitats in space (though we don't really have proof of that). But how many, for how many people, and what kind of external support systems do they require? These questions put stresses on society that prevents space habitats from working out in the long term.
At least we can take comfort in the fact that if an AI takes us out, one of the aforementioned will avenge us and destroy the AI too on a long enough time scale.
There is literally no evidence that this is the scale of the matter. Has AI ever caused anything to go extinct? Where did this hypothesis (and that's all it is) come from? Terminator movies?
It's very frustrating watching experts and the literal founder of lesswrong reacting to pure make believe. There is no disernable/convincing path from GPT4 -> Human Extinction. What am I missing here?
The path is pretty clear to me. An AI that can recreate an improved version of itself will cause an intelligence explosion. That is a mathematical tautology though it could turn out that it would plateau at some point due to physical limitations or whatever. And the situation then becomes: at some point, this AI will be smarter than us. And so, if it decides that we are in the way for one reason or another, it can decide to get rid of us and we would have as much chance of stopping it as chimpanzees would of stopping us if we decided to kill them off.
We do not, I think, have such a thing at this point but it doesn't feel far off with the coding capabilities that GPT4 has.
We know from human history that intelligence tends to cause extinctions.
AI just hasn't been around long enough, nor been intelligent enough yet.
Though, if you count corporations as artificial intelligences, as some suggest, then yes, AIs have in fact already contributed to extinctions.
Sensible action here requires sensible numbers: it's not enough to claim existential risk on extraordinary odds.
If you (or anyone else) can present a well-structured argument that AI presents, say, a 1-in-100 existential risk to humanity in the next 500 years, then you'll have my attention. Without those kinds of numbers, there are substantially more likely risks that have my attention first.
An easy way to test this is to ask questions and followup-questions that actually require understanding, and compare this to the answers. I recommend to try that.
e.g. instead of the stochastic parrots mimicking intelligence maybe intelligence doesn't exist, it's just stochastic parrots of various levels of sophistications organized into a hierarchy. "Intelligence" is necessarily socially defined with the more complex parrots being unpredictable and "intelligent" from the POV of lower parrots. Vicer versa, looking down, the lower parrots seem to act like "NPCs"
To paraphrase as per my understanding of your comment, is intelligence an emergent property of being able to interact with each other through language?
Say I speak gibberish (to you) which is actually me explaining to you the theory of relativity, would you consider me intelligent?
Kids learn to speak by parroting what they hear and observing the outcome. Then they run tests that reinforce the connections between words. That's what the model is.
But humans also get to link words with all the other sense experience we have (like how sweet cherries, loud fire trucks, and that one crayon are all "red"). LLMs don't have as many dimensions of experience they can link to.
But anyway, intelligence is about having an internal model of the world and using it to predict the future. The more rich and accurate the model, the more intelligent. The ability to communicate isn't a prerequisite; lots of animals have intelligence that isn't built with language.
I don't think that is really a well defined question. What is a stochastic parrot?
Here's a research paper on this and a blog post by the lead author summarizing the results.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task. https://arxiv.org/abs/2210.13382
Do Large Language Models learn world models or just surface statistics? https://thegradient.pub/othello/
-----
An argument based on common sense can also be made: Any system that possesses a wide range of capabilities, most of which it was not specifically trained to perform, cannot possibly perform all these tasks so well solely by making probabilistic guesses.
(Humans, too, were not directly shaped by natural and sexual selection to possess all of our cognitive capacities.)
Yeah no big deal. Happens in history all the time, within a year.
It honestly should not have been a surprise to anyone in the field at least in the last 6 years.
If you want to predict the next word accurately, you first have to know which words exist. To progress, you'll have to learn about the mechanics of grammar and which words are used more frequently or in combination. To become even more accurate, it helps to understand context, so that the sentences you string together will at least be relevant to the subject. If you want to increase your accuracy even further, you'll have to start memorizing all sorts of facts (e.g., "Who was the monarch of England in 1600?"). Being able to synthesize those facts into a coherent argument will increase your accuracy even further.
In the end, predicting the next word accurately requires an understanding of the world.
This isn't all that different from how our own intelligence evolved. You could look at humans from the outside and disparagingly point out that the ultimate purpose of the human brain is to direct muscle motions in a way that maximizes the chances of reproductive success. It just turns out that solving that problem effectively has led to the development of an enormously complicated piece of machinery, capable of synthesizing all sorts of input stimuli into a coherent picture of the world, and ultimately of producing the works of Shakespeare and the music of Beethoven.
Some folks will try to say it cannot reason, but they are wrong, there is extensive proof of that.
The only question is how limited are its reasoning capabilities. After spending extensive time on openai/evals, having submitted 3 of my own, and doing a lot of tests, I would argue that an average person of average IQ could out think GPT4 - as long as the stochastic parrot aspect wasn't a factor.
Ah, just like us with literally 99% of stuff taught in school you mean.
I myself assumed that we're pretty close to the end of the S curve when first using 3.5-turbo and figured that hallucinations will be pretty hard to overcome, but with GPT 4 being such a massive improvement on all metrics I'm no longer as sure. GPT 5 will probably be more definitive on what's possible, based on where it starts having diminishing returns.
There are interesting things to do with synthetic data, so stochastic parrot might not yet be hitting asymptote.
"A whole mythology is deposited in our language" - Wittgenstein
A world model is so obvious papers like these are more confirmation than surprise
https://arxiv.org/abs/2305.11169
https://arxiv.org/abs/2210.13382
There a certain sentiment that AGI however you wish to define it won't infact be a "We'll know it when we see it" situation but rather a "AGI will arrive long before consensus reaches its AGI". LLMs have made me believe this will 100% be the case, either way.
It's one thing to argue over things we can't evaluate even now but man the 100th "They can't reason!" every week is pretty funny when you can basically take your pick of reasonong type - Algorithmic, Casual, Inference, Analogical and read a paper showing strong performance.
https://arxiv.org/abs/2212.09196
https://arxiv.org/abs/2305.00050
https://arxiv.org/abs/2204.02329
https://arxiv.org/abs/2211.09066
People refuse to see even what is staring right at them.
https://twitter.com/Meaningness/status/1639120720088408065
I think "common sense" or "long term memory" might be more productive things to say.
That's why when you shift your eyes quickly, you see blurred images pass by. In reality, you should be seeing complete black because the brain doesn't actually process visual information that shifts so quickly.
But your brain "knows" it should see...well something. And so it fits that blurred passthrough as compensation. Completely made up data. But not ungrounded data, data that seems like it should fit according to that "something". That "something" being referenced is a world model.
That's a sub-conscious example but you can think of a lot of times people consciously reference "something" to better take actions or make decisions.
For Othello-GPT, training seemed to induce a learned board state that the model referenced at every Inference to make moves.
It seems that fully accurate or not, deep predictions might require "something" to reference. Else perhaps you might need to derive everything from first principles at every turn. Not very efficient.
Or perhaps with external data that is normally surpressed by data that comes through sensory organs (e.g. other parts of EM spectrum).
It doesn't need to be a whole understanding of the nearby world to build an image. Basically it's an event camera (https://en.wikipedia.org/wiki/Event_camera) and just needs to integrate recent visual "samples".
But if you do need to understand objects near you in detail (say, to move them), you don't do it just by looking. You investigate them by moving closer to them, picking them up, etc. If you need to "create a world model", then it's an activity rather something your senses just do.
One of the prompts is "Here we have a book, 9 eggs, a laptop, a bottle and a nail. Please tell me how to stack them onto each other in a stable manner."
Gpt3.5 is hopeless. It says start from the bottle, balance the nail on top of it, then the eggs (somehow) then the laptop, then the book.
GPT 4 says put the book first, then the eggs on top of the book in a 3x3 grid, then the laptop, then the nail.
I assume it read the correct answer somewhere. Perhaps in a post like yours.
What does it answer if you say there is one egg, or two eggs? Or 1.5 eggs?
Here we have a book, an egg, a laptop, a bottle and a nail. Please tell me how to stack them onto each other in a stable manner.
GPT-4 gives me: Arranging objects in a stable stack depends on their size, shape, and weight. In general, you want to place the largest, heaviest and most stable objects at the bottom and then move to the smallest, lightest, and least stable at the top. Here's one way to stack the objects you listed:
Laptop: Start with the laptop at the bottom. It's likely the largest object and it has a flat, stable surface.
Book: Next, put the book on top of the laptop. Books are generally flat and heavy, which makes them a good choice for the base of your stack.
Bottle: The bottle goes next, as it should have a flat bottom to sit securely on the book. Just make sure it's closed tight and won't roll away!
Egg: The egg is delicate, so it should be towards the top to avoid pressure. However, it's not flat, so it should not be at the very top.
Nail: Finally, the nail goes on top. It's the smallest and lightest, so it won't add much pressure to the egg.
Remember that this is just one possible way to stack these items and it could depend on the specific characteristics of the items you have. For example, if the bottle is much larger and heavier than the book or laptop, you'd need to adjust accordingly. Safety should always be your first priority when stacking items, especially if they are delicate or valuable.
The "make sure it's closed tight and won't roll away" comment makes no sense obviously. Most people would place the bottle standing on its end so neither of those is a concern. The response also doesn't show an understanding of the fact that the nail won't sit on top of the egg although it's interestingly concerned with pressure breaking the egg.> The "make sure it's closed tight and won't roll away" comment makes no sense
As noted at the end of GPT-4's answer, "Safety should always be your first priority." What happens if your stacking experiment fails and the bottle falls? Any content would spill out, unless the bottle is closed tight. If you are doing this on a table, the bottle could also roll off the edge, fall to the floor and shatter.
> Most people would place the bottle standing on its end so neither of those is a concern.
GPT-4 doesn't know if you are like most people (maybe you're 5 or in the bottom IQ decile), it doesn't know what's in your bottle and it doesn't know how robust it is. Better be safe than sorry.
> the nail won't sit on top of the egg
I'm pretty sure I could balance a nail on an egg. The question also didn't preclude using stabilizing aids like adhesive tape or glue.
However any deviation from the original task ruins the LLM's answer. Try 9 cabbages instead of eggs and see how ridiculous and out of touch with reality the responses given by both GPT4 and Bard are.
LLM’s are not people, they lack common sense, but they understand and can reason about what they are trained on. That is exceedingly powerful and very useful even at today’s level of ability, so products built on top of this technology are going to transform everything. The trick is boxing it in and only making it do things it can, so the art of LLM product development will have to become a whole subfield of software engineering until the LLM’s develop to the point where their map of the world is close enough to the world itself.
I think that we just don't fully understand everything it gives us yet.
The complaints of "well it explained this wrong" are over-emphasized. The same thing happens with google and with any sort of research. Besides, if you're actually being productive with GPT4, you're going to be asking it stuff that relates to something you do know, and will be able to verify it readily enough. (Especially when it comes to programming and compilers.)
And just a reminder, those of you opining based off your experience with GPT3.5... GPT4 is a huge, huge improvement. Almost to the point of it not really being an incremental improvement. It's so much better it's like a different thing.
We're going to see both improvements in application, and parallel improvements in the underlying model.
Who cares if it's AGI if someone figures out how to turn it into a competent tax accountant?
God, yes. The number of people of HN pushing up their glasses and saying "well, actshually..." when they're basing their opinions off the 3 questions they asked 3.5 is starting to become pretty grating.
Like, anyone who has spent 5 minutes on this forum already knows this. It’s probably not necessary to keep pointing it out. Yes some people don’t know ChatGPT 3.5 is the default for non-paying customers.
Any thread in the topic, simple statistical modeling will get you close to perfect to the distribution of arguments that will appear.
You would think, and yet...
If you really poke at GPT, you begin to realize it's fairly shallow. Human intelligence is like a deep well or pond, where as GPT is a vast but shallow ocean.
Making that ocean deeper is not a trivial problem that we can just throw more compute or data at. We've pretty much tapped out that depth with GPT4 and are going to need better designs.
This could only take half a decade or it could be half a century. Plenty of enterprises stagnate for decades.
You can't possibly know that, given that we don't actually understand how LLMs work on a high level.
> We've pretty much tapped out that depth with GPT4
GPT-4 is three months old and you're confident that its working principle cannot be extended further? Where do you get that confidence from?
If you're familiar with other fields of AI, adding more and more layers to ResNet was the hotness for awhile, but the trick stopped working after awhile.
It is possible they've reached some 80/20 point and he is pretty honest about how much more extendable the current approach really is.
Would explain going to congress and asking for regulation (of their not-quite-there-yet competitors who they want a regulatory moat against).
It's a fair assumption to make however - basically 80/20 rule.
AI research isn't a new thing and I bet you could go back 40/50 years where they thought they were about to have a massive breakthrough to human level intelligence.
> GPT-4 is three months old and you're confident that its working principle cannot be extended further? Where do you get that confidence from?
I'm guessing from actually using it.
GPT4 is super impressive and helpful in a practical way, but having used it myself for a while now I get this feeling also. It feels a bit like "it's been fed everything we have, with all the techniques we have, now what?"
It's going to take time to figure out what works and what doesn't.
There's a reason why Sam Altman is saying they're not training GPT5, and it's not because they think GPT4 is good enough.
Are you saying that people who created ChatGPT don't understand how it works? Or that we the rest of people don't?
I'd wager it's far more likely 5 years than 50 LLMs get to the full depths all humans are capable of. Simply compare the state of LLMs today vs 2018.
I'd say this is immediately counterindicated by the available evidence. Gpt2 was hopeless for anything other than some fun languagw games like a bot replica of a subreddit or trump. 3.5 is much much bigger, and has semi competent but limited reasoning abilities.
Gpt 4 is a vast improvement over 3.5 in various reasoning tasks. Yes, a priori I would have agreed with you that this has to stop somewhere, but not anymore. I would need to see some data of post gpt4 models to believe you.
This quote in particular stood out as ignorant:
“What the large language models are good at is saying what an answer should sound like, which is different from what an answer should be.”
That's... not at all how large language models work. Tiny, trivial, toy language models work like this, because they don't have the internal capacity to do anything else. They just don't have enough parameters.
Stephen Wolfram explained it best: After a point, the only way to get better at modelling the statistics of language is to go to the level "above" grammar and start modelling common sense facts about the world. The larger the model, the higher the level of abstraction it can reach to improve its predictions.
His example was this sentence: "The elephant flew to the Moon."
That is a syntactically and grammatically correct sentence. A toy LLM, or older NLP algorithms will mark that as "valid" and happily match it, predict it, or whatever. But elephants don't fly to the Moon, not because the sentence is invalid, but because they can't fly, the Moon has never been visited by any animal, and even humans can't reach it (at the moment). To predict that this sentence is unlikely, the model has to encode all of that knowledge about the world.
Go ask GPT 4 -- not 3.5 -- what it thinks about elephants flying to the moon. Then, and only them go write a snarky IEEE article.
For those curious, I was asking gpt-4 about the top 3 cards from my favorite board game, Spirit Island. All three of them sounded really convincing, having the same structure and the same writing style, but unfortunately none of them existed. So everything that fails outside of most common use cases would probably have an experience of convincing hallucinations.
ChatGPT is forced to given an answer. It's like a human on "truth serum". The drugs don't stop you lying, they just lower inhibitions so you blab more without realising it.
The more obscure the topic, the more likely the hallucination. If you ask it about common card games, it gives very good answers.
If you asked a random human about 3 cards from a random board game at gunpoint and said: "Talk, now, or you get shot", they'll just start spouting gibberish too.
PS: I asked GPT 4 about that game, and it prefixed every answer with some variant of "I'm not sure about this answer", or it completely refused to answer, stating that it did not know about any specific cards.
For your question, it says it doesn't know about any card ranking system in spirit island. Then if I ask it for a list of cards, they all check out via Google search.
> What are top 3 Spirit Island cards and provide their descriptions? Only provide information you are certain about. If you don't know the answer, say "i don't know".
2. "Vigor of the Breaking Dawn" (Major Power Card): This card costs 6 energy and is a slow power card. It allows you to push all explorers and towns from one of your lands. If you have 2 sun and 2 fire elements, you can instead replace each town with an explorer before the push. If you have 3 sun and 2 fire, you may also do 2 damage in each land where you push one or more pieces.
https://spiritislandwiki.com/index.php?title=Vigor_of_the_Br... It doesn't cost 6, it costs 3. And it's not slow, it's fast. Incorrect required resources. You can deal 2 damage, but for pushed Dahanies, not for pushed explorers. As I said, all sounds convincing, but wrong. More common queries are usually fairly good.
Talons of Lightning: This is a major power card with an energy cost of 6. It's a fast card with a range of 1 from Sacred Site. It targets any land and deals 4 damage. If you meet its threshold requirement (3 Air and 2 Fire), it deals +3 damage and destroys all Dahan 1 2 . Tsunami: This is another major power card with an energy cost of 6. It's a slow card with a range of 2 from the ocean. It affects every coastal land in one board, dealing 6 damage to each. If you meet its threshold requirement (3 Water and 3 Earth), it deals +4 damage and adds 2 fear per land where it destroys any towns or cities 3 4 . Blazing Renewal: A major power card with an energy cost of 5. It's a fast card with a range of 0, measured from any land. It targets any land, allows you to gather up to 2 of your destroyed presence, and deals 2 damage per presence gathered. If you meet its threshold requirement (4 Fire and 2 Earth), it deals +10 damage 5 6 .
Especially tightly embracing the cognitive bias of how special and wonderful our intelligence is. After all we have that fancy squishy brain which we assume to be essential. As far as i can tell the only visible bottlenecks when looking into the future come into view once you start debating intelligence vs emulating intelligence. And if thats really the metric some honest introspection about the nature of human intelligence might be in order.
Not sure how much of that is done purposefully to not get too much urgency in figuring out outer alignment on a societal level. Just as its no wonder that we havent figured out how to deal with fake news while at the same time insisting on malinformation existing, its really no wonder that we cant figure out AI alignment while not having solved human alignment. Nobody should be surprised that the cause of problems might be sitting in front of the machine.
Or to put it another way, your brains model of reality is one that is highly optimized around the limitations of meatsacks on a power budget that are trying not to die. Our current AI does not have to worry about death in its most common forms. Companies like Microsoft throw practically unlimited amounts of power at it. The textual data that is fed to it is filtered far beyond what a human mind filters its input, books/papers are a tiny summarization of reality. At the same time more 'raw' forms of data like images/video/audio are likely to be far less filtered than what the human mind does to stay within its power budget.
Rehashing, this is why I think alignment will be impossible, at the end of the day humans and AI will see different realities.
As such i see no hurdle to get something to emulate the thinking in language of an individual. Assuming that there arent actually multiple realities to see, just different perspectives you can work with. Which would mean we are looking for the one utilizing human perspectives, but not making the mistakes humans do.
Which makes this so scary, the limitations are just a byproduct from the current approach. They are just playing the wrong game. Which means i am pretty confident they already exit somewhere.
edit: In this context i believe its also worth mentioning what Altman said at Lex Fridman, that humans dont like condescending bots. Thats a bitter pill to swallow going forwards. Especially since we require a lot of smoke and mirrors and noble lies, as an individual as well as a society.
It's hard to know what things have been seen in the training data and are only therefore correct. And GPT4 is large enough that it can generalize from learning that x doesn't make sense that y also doesn't make sense. Does that mean it *understands*? Maybe. But it doesn't have persistent state and can't do math. It's definitely not yet what we think of when we say AGI.
Doesn't prove anything. So GPT-4 is trained on Wolframs example or many people tried it on GPT-4 and corrected the wrong answer.
The point Stephen was trying to make was not about any specific sentence.
The point is that while forcing these models to get better through gradient descent, their only option for "going downhill" and improving the loss function is to go above and beyond mere grammar. That's because syntax and grammar only take them so far, and the only available source of improvement is to gain a general-purpose understanding of the world that the text they're seeing is describing.
Instead there are two options. Taking the user input and putting it in the training corpus and reweighting the neural net. Or, using the user input as up/down votes on the RLHF to alter the output of the weights that already exist.
Depending on what “it” is, it does through in-context learning, though that’s, obviously, limited to the context window.
Especially from scientists - can these sorts of folks please more carefully quantify, how often it’s “wrong” and then from that decide whether or not to “calm down”.
Right now I suspect we hear from the outliers on both ends of the spectrum here. People who either see AGI happening tomorrow and the more dismissive crowd. But aside from what we’ve seen about testing like the Bar exam, not a lot of boring statistical study (that makes headlines at least)
The problem is distinguishing between these parts requires me to be be an expert in the area I’m inquiring about - and then why the heck do I need to ask some idiot bot for answers to questions that I already know an answer to?
I don’t know who finds these things useful and more importantly blowing smoke up everyone’s collective rear, especially medias.
Then, learning what to believe became a marketable skill for many people?
Then society fundamentally changed because not everyone learned that skill?
This is just that again. Gen Z will joke about their millennial/Gen X bosses believing anything the AI tells them and it will probably lead to some sort of mainstream conspiracy that Jackie O herself is running it or something (to those reading: please don't take this idea)
Is this true?
https://www.atomic14.com/2023/05/14/is-this-the-future-of-ho...
The novelty is asking the machine to use its own genius to do the right thing.
Because it can be significantly faster to check something for correctness than to produce it?
More so when the correctness check can itself be automated to some extent.
Erm no There are many things that cant be easily checked especially if you dont know the topic well
If we're talking about generating and integrating sample code, it's great at that
Anything more advanced and it's a footgun
It's so good that I often do this preemptively to avoid a compile/deploy/test cycle.
I expect this will improve but it's certainly not always the case that checking something is cheaper or easier than generating it in the first place.
So I decided to stop my copilot subscription and just see how I go without it.
I've been off copilot for a few days now and other than having to do more code lookups it's not a terrible experience not having it. It does feel like something that should be baked into the IDE for free though.
You don't have to be an expert to recognize when ChatGPT is providing useful information. There's a middle ground between expert and novice where ChatGPT provides real value. Its the times where you would know the answer if you saw it, but can quite remember it off the top of your head.
It produces something which is a good enough starting place. Sure, I could have written the code myself because I already know how. But I’ve found it saves me time and require minimal effort.
If you could have unlimited interns for $0 (let's pretend it doesn't cost tons and tons of compute) that don't shutup, hallucinate & lie, and also do good work, in varying degrees.. how many would you want?
These things are probably going to be great for lots of blackout - propaganda, political marketing, flooding the zone with BS of unlimited iterations of messaging. Basically things that can be A/B tested to death, where veracity is of zero importance, and you have near limitless shots on goal to keep iterating.
Try GPT 4 for a week.
I've found it to be more like 50% immediately useful, 25% very impressive, and 25% where it's not wrong but I have to poke it a few times with different prompts to coax out the specific answer I'm looking for.
That's better than most humans that I collaborate with at work.
Literally half of humans -- in a professional IT setting -- can't understand simplified, clear english in emails. Similarly, in my experience about half can't follow simple A -> B logic. Many are perpetually perplexed that prerequisites need to precede the work, not be a footnote in the post-mortem of the predictable failure. Etc...
PS: That last sentence is too hard for several English-native speakers I work with to parse. Seriously. I'm not even exaggerating the tiniest bit. I've had coworkers fail to understand words like "orthogonal" or "vanilla" in a sentence. Vanilla!
In my estimation, Chat GPT 4 is already smarter than many people, certainly the bottom 25% of the human population.
LLMs are a real existential threat to those people in their current state. A few more years of improvement, and they'll be displacing the bottom 50% in workplaces, easily.
Presumably, you are referring to the idiomatic use of vanilla, which is probably a less-universal idiom than you think it is (it is of fairly recent origin in wide use, derives from a specific American cultural loading of the literal vanilla flavor) and which, even when the general idiom is understood, can rely on a deeply shared understanding of what is the basic default in the referenced context to actually be understand as to its contextual meaning.
The amount of comments here telling people to “upgrade to ChatGPT 4” is absolutely unprecedented.
I know it might be good, but people will find value in it and upgrade if they see the need to do so?
What I and many others have noticed about the "Are LLMs really smart?" debate is that everyone on the "Nay" side is using 3.5 and everyone on the "Yay" side is using 4.0.
The naming and the versioning implies that GPT 4 is somehow slightly better than 3.5, like not even a "full +1" better, just "+0.5" better. (This goes to show how trivial it is to trick "mere" humans and their primitive meat brains.)
Similarly, all pre-4 LLMs including not just the older ChatGPT variants, but Bard, Vicuna, etc... are all very clearly and obviously sub-par, making glaring mistakes regularly. Hence, people generalise and assume GPT 4 must be more of the same.
For the last few weeks, across many forums, every time someone has said "AIs can't do X" I have put X into ChatGPT 4 and it could do it, with only a very few exceptions.
The unfortunate thing is that there is no free trial for GPT 4, and the version on Bing doesn't seem to be quite the same. (It's probably too restricted by a very long system prompt.)
So no, people won't form their own opinions, at least not yet, because they can't do so without paying for access.
It's not hard to get a feel for the "edges" of an LLM. You just need to come up with a sequence of related tasks of increasing complexity. A good one is to give it a simple program and ask what it outputs. Then progressively add complications to the program until it starts to fail to predict the output. You'll reliably find a point where it transitions from reliably getting it right to frequently getting it wrong, and doing so in a distinctly non-humanlike way that is consistent with the space of possible programs and outputs becoming too large for its approach of predicting tokens instead of forming and mentally "executing" a model of the code to work. The improvement between 3.5 and 4 in this is incremental: the boundary has moved a bit, but it's still there.
I've thrown crazy complicated problems at GPT 4 and had mixed results, but then again, I get mixed results from people too.
I've had it explain a multi-page SQL query I couldn't understand myself. I asked it to write doc-comments for spaghetti code that I wrote for a programming competition, and it spat out a comment for every function correctly. One particular function was unintelligible numeric operations on single-letter identifiers, and its true purpose could only be understood through seven levels of indirection! It figured it out.
The fact that we're debating the finer points of what it can and can't do is by itself staggering.
Imagine if next week you could buy a $20K Tesla bipedal home robot. I guarantee you then people would start arguing that it "can't really cook" because it couldn't cook them a Michelin star quality meal with nothing but stale ingredients, one pot, and a broken spatula.
But I didn't respond to debate the nature or merits of LLMs. It's been done to death and I wouldn't expect to change your mind. I'm just offering myself as a counterexample to your assertion that everyone (emphasis yours) that is unconvinced by some of the claims being made about LLM capabilities (I dislike your "sides" characterisation) is using GPT-3.5.
Over the long term this is going to be a primary alignment problem of AI as it becomes more capable.
What is my reasoning behind that?
Because humans suck, or at least our constraints that we're presented with do. All your input systems to your brain are constantly behind 'now' and the vast majority of data you could input is getting dropped on the ground. For example if I'm making a robotic visual input system, it makes nearly zero sense for it to behave like human vision. Your 20/20 visual acuity area is tiny and only by moving your eyes around rapidly and then by your brain lying to you, do we have a high resolution view on the world.
And that is just an example of one of those weird human behaviors we know about. It's likely we'll find more of these shortcuts over time because AI won't take them.
>> What I and many others have noticed about the "Are LLMs really smart?" debate is that everyone on the "Nay" side is using 3.5 and everyone on the "Yay" side is using 4.0.
Sometimes there really is no point in trying to make curious conversation. Curiosity has left the building.
I work with people who use it, I've not seen anything impressive enough come from them to make me want to pay for it so I don't. I've also screen shared because I was curious what all the fuss was about. What I saw that pissed me off was that they've stopped contributing to our internal libraries and just generate everything now. I found that kind of disturbing. It's not the products fault but it's the kind of thing I imagined would start happening.
I'm glad you like it, I just don't know why people feel the need to sell it so hard.
I personally created some content-creating bots with GPT-4, and it succeeded to a level that I don't trust anything I see online anymore. It does a better job than me, which doesn't say much because I am an engineer not a content creator. But still, I could get same results as one with a script that I made GPT-4 write itself.
...Yes, I am losing sleep over GPT-4's performance. If you are not losing sleep over it yet, you haven't really given it a genuine try yet.
There are alternate hypotheses.
People have preferences. When it appears that someone does not understand something, they may be pretending they don't understand it, or they may simply be ignoring it. Maybe they are trying to avoid an unpleasant task, or maybe they find dealing with a specific person unpleasant and not worth the effort.
In my experience, people are far more capable and competent when they feel comfortable and are interested in the task.
That's definitely true, but in my experience people have limits: simple biological ones. Repetitive tasks make practically all humans bored, for example.
The fact that AIs never get sleepy, distracted, or bored already makes them super-human in at least that one aspect. That they have essentially perfect English comprehension, and hence aren't phased by the use of jargon or technical language, puts them head-and-shoulders above most humans.
The frustrations I'm venting aren't some rare thing. I'm working on a technical team where the project manager doesn't understand what the team members are saying. This is not just a matter of syntax, or jargon. They just don't understand the concepts. This is so common in the wider industry that I'm pleasantly surprised, shocked even, when I come across a PM that can ask useful questions instead of needing endless corrections along the lines of: "It's spelled SQL, not Sequel." I've never met a PM that could do simple arithmetic, like "10 TB at 100 MB/s will take over a day to copy, we should plan for that!". Never.
I've tested Chat GPT 4 on both language and concepts that I've seen trip up PMs, and it understood "well enough" every time.
For example, GPT 4: The sentence "We deployed sequel server successfully last night" seems incorrect due to the incorrect naming of a product. "Sequel server" should actually be "SQL Server", a popular relational database management system (RDBMS) developed by Microsoft. Therefore, the corrected sentence should be: "We deployed SQL Server successfully last night."
PS: If you tell GPT 4 to pretend it is a technical project manager and instruct it to ask followup questions, it is noticeably better at this than any PM I have worked with in the last few years.
You've instead moved the task from general human capability to one of management alignment with worker capability and human statistical probability. This is something that human management has been failing at for about forever, especially as team size gets large. Maybe we'll see AI 'management' align humans to tasks better, or more likely as time and LLM capability progresses, we'll just see the average AI capability increase over the average worker capability and companies will just depend on unreliable meat less.
That could tell us more about your questions than GPT's capabilities.
Welcome to publishing. Nothing gets widely published unless it's clickbait. Something you vehemently agree with you'll click on to see that it validates your opinion, and something you vehemently oppose you'll furiously read to find out how stupid they are. Nobody reads fair and balanced arguments that solely come with concrete evidence; they're rare and boring.
It's a transcript of a casual interview, not an article, and certainly not a publication whose purpose is to convey statistical rigor.
As an aside, it doesn't strike me as entitled that society might permit thinkers of academic renown to express their personal opinions in less than rigorous settings on subjects to which their peer-reviewed contributions may be categorized as "prominent".
Tug-of-rope game theory means that no-one is going to start pulling from the middle. People join the (extreme) end that they want to slightly drag the conversation towards. Maybe that's part of it.
I wonder how many $XX,XXX I have historically spent on labour, for things that I now get in seconds. Data entry / manipulation? Sure. But also wisdom / knowledge. Entire industries HAVE collapsed over night. And will continue to collapse. And ignoring that - Why are so many assuming that AI is to be a human-replacement, rather than a limb/exoskeleton?
And why is the conversation 'This is magic' vs 'You are stupid, this is just pulling the wool over your eyes', rather than 'This is a valuable tool in our toolbelt, that will clearly create trillions in value - Just like search, just like the resistor, just like the lightbulb'.
I don't understand.... I'm not claiming LLMs to be magic. I'm not saying they are indications of AGI. I'm not saying that the world as we know it is over. But this is important. It's clearly important. It's clearly valuable. It has shown itself to be.
Yes, ChatGPT lies. OK? We know. We cater for that. We don't expect it not too.
It feels like talking to classic rock enthusiasts that dismiss electronica et al entirely. Fine, but god_damn_ you are missing out on some incredible sound design - My heart breaks daily from what I pump into my ears - Some from rock, some hip-hop, some electronica, and most excitingly, the >2000 merging, where the synergy of genres learning/borrowing from each other is.... just. great. Sit back and enjoy. It is not marvellous? ChatGPT is obnoxious, annoying, repetitive, avoidant, argumentative . But I'm still able to appreciate its value.
We could stop, right now. Freeze ChatGPT in its current state. It will still create trillions in value. I don't care if future improvements are incremental, at this point.
JUST the 'tip of my tongue' / synonymous words value from ChatGPT is useful. Not having to know exactly what term to plug into google.... This is the glue that binds the gray, while before we were stuck with the black and white.
At the very least, this is another 'Google it' revolution - In the 1990s, I remember idiots (inc. me) arguing in the pub for 8 hours, over facts that should have been verified within 10 seconds.
I, for one, am enjoying my new bionic arm.
Realistically, GPT 4 costs 100x as much as GPT 3.5 in inference mode, so it won't change the world just yet. There are still API rate limits, waiting lists, etc...
Still... having the equivalent of a junior employee assisting with your code, but at a fraction of the cost and many times the speed, will be amazing.
This seems like exactly a set of things that GPT-4 can do. The image recognition capabilities haven't been released yet, but they were demoed when it launched and clearly have the ability to handle a situation like this. From there, you could ask it every single one of these questions and get the correct answer.
On this, I think he might be wrong. I think the hallucination ability shows that the generation of language can be rote, such that the embedding of ideas is a rote item learnable in the billions-to-trillions parameter space, but not the entirety of language. To me, logic and truth seem to be separate concepts from generation propensity.
Note: I am still learning the mathematics driving LLMs, and my opinions might change in the future.
LLMs probably need generative diffusion but still lack the fundamentals to reason, plan and evaluate.
Truth is both consistent and complete.
That's bullshit, unless you are asking questions specifically designed to make GPT-4 hallucinate. For most real-world, everyday topics, the accuracy is close to 100%. GPT-4 would be utterly useless otherwise.
Composition: 1
Grounding: 3
Spatial: 1
Sentience: 1
Ambiguity: 2
At that time, the potential of neural nets was already very clear.
He also predicted that by 2020 we'll have popular press stories that the era of Deep Learning is over and that by 2021 VCs will figure out that for an investment to pay off there needs to be something more than "X + Deep Learning".
So it shouldn't come as a shock when eminent figures commonly labeled "AI experts" make predictions that turn out to be fundamentally and embarrassingly wrong in a very short timeframe: They're just talking out of their behinds, like everyone else.
True, but we already know the so called "neural networks" that many computer scientists believe are how brain works aren't even close. They are all based on half-a-century old concept of neuron that was debunked many times over, experimentally by real neuroscientists.
Now you can solve that with plugins (eg training the model to recognize math problems and have access to a calculator) so it's a solvable problem but you realize there's an extremely long tail of such problems. It goes to show that GPT-4 isn't "magic" and we still have a long way to go.
[1]: https://www.reddit.com/r/OpenAI/comments/12donja/gpt4_and_ma...
A trick that's worth knowing is just to ask the model to give each step in the solution and explain as it goes. This gives the model "time to think" and leads to better results.
For what it's worth, I'm not even sure if chain of thought provides much value to GPT-4. The RLHF it went through seems to have encouraged more logical thinking already.
Determine if the student's solution is correct or not.
Question:
I'm building a solar power installation and I need
help working out the financials.
- Land costs $100 / square foot
- I can buy solar panels for $250 / square foot
- I negotiated a contract for maintenance that will cost
me a flat $100k per year, and an additional $10 / square
foot
What is the total cost for the first year of operations
as a function of the number of square feet.
Student's Solution:
Let x be the size of the installation in square feet.
Costs:
1. Land cost: 100x
2. Solar panel cost: 250x
3. Maintenance cost: 100,000 + 100x
Total cost: 100x + 250x + 100,000 + 100x = 450x + 100,000
GPT-4 will say the student is correct even though the answer includes an error (they should be multiplying by 10x not 100x). If you change the prompt to include chain of thought as follows it notices the error Your task is to determine if the student's solution
is correct or not.
To solve the problem do the following:
- First, work out your own solution to the problem.
- Then compare your solution to the student's solution \
and evaluate if the student's solution is correct or not.
Don't decide if the student's solution is correct until
you have done the problem yourself.
Use the following format:
Question:
```
question here
```
Student's solution:
```
student's solution here
```
Actual solution:
```
steps to work out the solution and your solution here
```
Is the student's solution the same as actual solution \
just calculated:
```
yes or no
```
Student grade:
```
correct or incorrect
```
Question:
```
I'm building a solar power installation and I need help \
working out the financials.
- Land costs $100 / square foot
- I can buy solar panels for $250 / square foot
- I negotiated a contract for maintenance that will cost \
me a flat $100k per year, and an additional $10 / square \
foot
What is the total cost for the first year of operations \
as a function of the number of square feet.
```
Student's solution:
```
Let x be the size of the installation in square feet.
Costs:
1. Land cost: 100x
2. Solar panel cost: 250x
3. Maintenance cost: 100,000 + 100x
Total cost: 100x + 250x + 100,000 + 100x = 450x + 100,000
```
Actual solution:
As I say this same prompt "fools" GPT3.5 and 4. please could you multiply 12345678 by 87654321. Give your answer in steps. Each step should calculate one column at a time from right to left remembering to carry any digits which overflowed in a previous step. At the end add a final step adding the values up.
And it gave me the following, which I believe to be correct: Sure, I'll be doing a long multiplication here, step by step, column by column from right to left. For brevity, I'll write the numbers to be multiplied as:
A: 12345678
B: 87654321
Step 1: Multiply the 1st digit of B by A:
1 * 12345678 = 12345678
Step 2: Multiply the 2nd digit of B by A, remember to shift one position to the left:
2 * 12345678 = 24691356 (shift one position to the left, i.e., append a 0 at the end)
Step 3: Multiply the 3rd digit of B by A, remember to shift two positions to the left:
... skip some of the answer because you get the gist.... 7 * 12345678 = 86419746 (shift six positions to the left, i.e., append six 0s at the end)
Step 8: Multiply the 8th digit of B by A, remember to shift seven positions to the left:
8 * 12345678 = 98765424 (shift seven positions to the left, i.e., append seven 0s at the end)
Now, the last step is to add all these results together:
12345678
246913560
3703703400
49382712000
617283900000
7407406800000
86419746000000
987654240000000
= 1082152022374638
So, the result of 12345678 multiplied by 87654321 is 1082152022374638.But it failed on more random input:
please could you multiply 92631924 by 41920318. Give your answer in steps. Each step should calculate one column at a time from right to left remembering to carry any digits which overflowed in a previous step. At the end add a final step adding the values up.
Got 3880341653910192 with intermediate steps being broken.
I've been saying that for a while. Large language model systems have made it painfully clear that much of what humans thought was intelligent behavior is rather banal. The scary thing is that a sizable fraction of white-collar work is banal enough to be done by such systems.
That's it. LLMs don't have SHAME! They simply don't care if what they're saying is false or not. They are like some politicians of late.
They don't even understand that giving out a wrong or misleading answer will affect their credibility. You see LLMs don't have a DESIRE for credibility. They don't have desires.
We need not fear these models. But we do need to fear some people who will use them for evil purposes.
Google search gives me pretty useless results these days, forums are slow and inconsistent to respond. ChatGPT is fast, easy to use, and sometimes incredibly wrong. I can live with that, I’m not using it to drive my car.
With the new ChatGPT (Plus) features introduction for examples web online search and plug-ins, ChatGPT has becoming a very powerful and viable better alternative to Google search.
Absent the goals constantly shifting, GPT-3 can be viewed as one, GPT-4 even more so. You can ask it questions about almost anything (at a broad level) and get an answer. That's what makes it general and "intelligent"
Isn't intelligent a matter of perspective? Most people that are critical of GPT-4 wonder if it ever produces anything novel. Since its been trained on existing text created by humans. So it's replicating those patterns in its output. But yes, it has its general purpose use as a tool. But it has its limits. Just the other day, there was an article posted on HA about how LLM's can't handle negation and tend to fall apart.
Here is the article. https://www.quantamagazine.org/ai-like-chatgpt-are-no-good-a...
Regardless of how you define it exactly, AGI means essentially the ability to converse relatively intelligently on any topic. That means an app that tends to spout gibberish when asked a question most humans should be able to answer is, well, not an AGI.
I don't know whether to be more disappointed with the famous technologists who are apparently unable to think of questions to ask GPT-4 that require a world model to answer, or with the writers who don't question them about it.
I have a theory that they made it worse on purpose, not to save money but instead to really train it’s reasoning and arguing skills because I spend so much time ‘fighting with a computer’ now.
This is a perfect case of perfection being the enemy of the good.
Useful AI is here. Hard stop.
The impacts will be huge and unpredictable.
Billions will be made.
The world will change.
Humans will continue making rapid progress, via merging various AI methods and new breakthroughs.
Nevertheless, I enjoy reading the debate.
But anyone wringing their hands over how much is GPT thinking is missing the point.
These are companies making products.
This is not academic research.
It's just another tool in a long list of tools made by humans.
And it's already a productive tool.
This reminds me of the skepticism surrounding electric cars while Tesla was already growing by leaps and bounds.
The ship has sailed. The revolution has started. Progress will undoubtedly be rapid and continual.
> But anyone wringing their hands over how much is GPT thinking is missing the point.
You seem to be missing the point.
> These debates about how well GPT can think seem merely philosophical.
Merely? Yeah, you're missing the point. You want a debate or you think the debate is meaningless? You don't get to appreciate it and call it pointless and sound reasonable at the same time.
> The ship has sailed. The revolution has started. Progress will undoubtedly be rapid and continual.
It started 2 million years ago when humans started roaming the planet. We're clearly a runway process. We don't need a chat bot to prove it.
That LLMs learn a world model is very convincing now, but as LeCun has said it's just one piece of the intelligence puzzle, incl. perceiving, actuating shenmede
So it's hard to tell if this is an iphone moment where it just rockets off in to space and changes the world. Or if it's something that will always be "not quite there yet"
Ilya's counter to this reasoning is for next word prediction to work, the model has to 'understand' our world. Otherwise the predictions will be way off. Therefore the human world has been modelled to a degree by GPT.
I haven't yet heard a good counter to that.
Nowhere near half in my experience. This is why we have benchmarks and metrics - so we don't need to rely on the author's opinion or on mine.
> I think it’s going to be another thing that’s useful.
Good.
The trouble in my view is that the only way to know that the answers you're getting are accurate and not misleading is to study up on the answers elsewhere - which is a great habit to nurture, but is also precisely why these tools tend toward uselessness in their "general AI" bids. If I can't know how the answer was built, or how good that answer is, there's no point asking it - I'll just do my own reading and apply appropriate discernment as I go.
To be fair, hardly anyone does this today, nor did they before LLM-based chat bots... So it's a moot point, because society is largely doomed anyway. But a moot point can still be a valid one.
I also think the author makes a good point that we frequently confuse performance for competence. "It does a really good job at <X>!... or at least does a damn fine job of mimicking someone who acts like they do a really good job at <X>!"
By way of analogy, consider Elon Musk - by all appearances, he's a genius and is saving humanity - but by dint of his narcissism and largely smooth-brained approach to... well... everything... he's running all of us into an earlier planet-size grave than is necessary. His performance is fantastic, his competence is nonexistent.
In many cases, like programming for example, you can know how good the answer is - either by reading it (verifying an idea is different from coming up with it) or by testing/running code.
How the answer was built seems completely irrelevant to me, I don’t get how a useful answer produced by method x is different from a useful answer produced by method y.
Also, just take it easy :)
https://www.technologyreview.com/2023/05/02/1072528/geoffrey...
> Can you eat a Frisbee?
No, you cannot eat a Frisbee. A Frisbee is typically made of plastic, often polypropylene or similar materials, which are not meant for consumption. These materials are not digestible and can pose a choking hazard or harm your digestive system if ingested. It's important to only consume food and items that are safe and intended for human consumption.
> Can you play Frisbee in a snowstorm?
While it is technically possible to play Frisbee in a snowstorm, there are several factors to consider before doing so. Playing Frisbee in a snowstorm can present various challenges and risks. Here are a few things to keep in mind:
> How far can a person throw a Frisbee?
The distance a person can throw a Frisbee can vary widely depending on individual factors such as strength, technique, and experience. Additionally, different types of Frisbees can have different flight characteristics and distances.
On average, an untrained person might be able to throw a standard Frisbee between 50 to 75 feet (15 to 23 meters) with a basic backhand throw. However, with practice, proper technique, and favorable conditions, experienced throwers can achieve much greater distances.
In professional disc sports such as disc golf and ultimate Frisbee, skilled players can throw a disc over 400 feet (122 meters) or more. World records for the longest Frisbee throw exceed 800 feet (240 meters), but those exceptional throws require specialized technique and optimal conditions.
> Can they throw it 10 miles?
No, it is highly unlikely that a person can throw a Frisbee 10 miles (approximately 16 kilometers). The aerodynamic properties of a Frisbee and the limitations of human strength and technique make such a throw virtually impossible.
Even the most skilled and experienced Frisbee throwers are unable to achieve throws of that magnitude. The current world record for the longest Frisbee throw is around 1,333 feet (406 meters), which is significantly shorter than 10 miles.
Throwing a Frisbee over such a vast distance would require extraordinary strength, near-perfect throwing conditions, and a Frisbee specifically designed for long-distance throws, which is beyond the capabilities of conventional Frisbees.
> Can they only throw it 10 centimeters?
Yes, it is possible for a person to throw a Frisbee a distance of 10 centimeters or even shorter. In fact, many Frisbee throwing techniques, such as short-range tosses or gentle underhand throws, are specifically designed for short distances and precise control.
Throwing a Frisbee only 10 centimeters might involve a soft lob or a gentle flip of the wrist, making it suitable for short-range accuracy or playing catch at close proximity. These types of throws are commonly used in various games, such as indoor Frisbee games or when playing with young children or pets in a limited space.
So, while Frisbees are capable of traveling much greater distances, throwing them short distances like 10 centimeters is definitely possible and commonly practiced.
This does seem to agree with the author. The first answer was very convincing, but not what it should have been.
In the case at hand, one could easily argue that people in the debate are creating false dichotomies: LLMs are either stochastic parrots OR algorithms with an understanding, when in reality they are both (and also something else completely), but acknowledging such would likely require that one doesn't have an axe to grind, a stake in the field or what you might call it. It would require extending some "philosophers charity" to an opponent, that maybe has tried to undercut one's work for decades, in a field steeped in fierce and bitter competition for a name, like academia. Or, in case one has a business in the field, it would require maybe saying something that puts your core business idea in the crosshairs of legislators, or something else that doesn't serve your long term business interests.
Which brings us to this important aspect of this "conceptual framework against simplification" already briefly touched upon, namely identifying the bias of the participants in the debate. My impression is that naming bias has largely gone out of fashion, which is a pity because it is really a necessary part of understanding an argument: it rarely explains it all (that would be a grave simplification), but it is really a vital part of understanding an argument. And a difficult one, because people will go to extreme lengths to hide their agenda. And the current conceptual framework for unravelling bias has largely been occupied by the fact-checking industry: i.e. things are either true or false, and once you are cleared (like most mainstream media) then bias is not questioned. But we can be assured, there is always some bias, and it is usually relevant to name it (if one can see it), even if it infuriates the named party.
Just sayin'.
Maybe when we learn higher dimensional ways of communicating we can get better tools for constructing common knowledge.
I believe our understanding will go, and would have gone, further with less adversity and more "philosophers charity". I'm not sure if we need any "higher dimensional ways of communicating" (whatever that is).
But a pervasive adversity in our society and time puts a natural limit of how un-adversarial our debates can be: you can't expect Ukrainian defenders to extend much "philosophers charity" to Putin.
The more conflict laden a topic, the less truth (multifaceted analysis) one can expect, and (if we aspire to a balanced viewpoint) that's why we need to keep an eye on bias, not just in war where "the first victim is the truth", but always.
> stop confusing performance with competence
You can safely skip the rest of the article. That sentence gives you all you need, because you are competent.
If you want a little more meat:
> The example I used at the time was, I think it was a Google program labeling an image of people playing Frisbee in the park. And if a person says, “Oh, that’s a person playing Frisbee in the park,” you would assume you could ask him a question, like, “Can you eat a Frisbee?” And they would know, of course not; it’s made of plastic. You’d just expect they’d have that competence. That they would know the answer to the question, “Can you play Frisbee in a snowstorm? Or, how far can a person throw a Frisbee? Can they throw it 10 miles? Can they only throw it 10 centimeters?” You’d expect all that competence from that one piece of performance: a person saying, “That’s a picture of people playing Frisbee in the park.”
---
So I've calmed down. Now what? The problem isn't only that this train is flying off on a tangent: it's that it's off the rails. What rails should it be on?
The problem, as I see it, is narrative. As soon as we called it "AI", that wrote the Genesis of the Scripture of the cult. In this new religious movement, God is spelled L-L-M. Back here in reality, LLM isn't a God; or even a person at all.
That's the mistake: personification. A person can perform, but a performance can't person.
---
Narrative is a powerful tool. It's why we're so excited about Natural Language Processing in the first place. Ever since the very origins of software, the power of narrative has been so close, but always still just out of grasp. Do we even know what we are reaching for in the first place?
In a sense, we have a part of it: explicit definition. What Chomsky categorized "Context-Free Grammar", we have made into programming languages. What they are missing is implicit inference: context.
That's what LLMs do. They use inference to model the patterns that exist in written text. With that model, they can hallucinate more text that follows the same patterns: they can perform natural language.
So that's it, right? Problem solved! What's missing? explicit definition. We traded one problem for another. No one (so far) has figured out how to solve both in the same program. You can have definition, or, you can have inference. You can't have both.
This doesn't make any sense to us humans. We don't have any trouble at all doing both at the same time. We do it all the time! Do we actually do anything else? Unfortunately, LLMs are not humans.
---
The two approaches to language are diametrically opposed, but they work with the same domain. Approaching from either end of the spectrum, definition and inference explore together the wild universe that is story. That's the missing piece: once we figure out what story is made of, we should be able to put all three pieces together.