There are extremely well researched worms with only a few hundred neurons which we cannot yet simulate with anything resembling accuracy. How can that statement be true, if LLMs are close to delivering superhuman intelligence?
There are extremely well researched worms with only a few hundred neurons which we cannot yet simulate with anything resembling accuracy. How can that statement be true, if LLMs are close to delivering superhuman intelligence?
LLMs are a huge step forward. Sure, they might not be the thing to ultimately deliver superhuman intelligence. But it's unfair to say that we're not any closer at all.
> We don't simulate a bird's bones and muscles
Isn't really true, is it?
We simulate the bones with an airframe, and we simulate the muscles with a prop/jet. The wings are similarly simulated.
By your definition of simulate, I think artificial neural networks are absolutely on their way to simulate intelligence
It's true that Markov chain generators have existed for years. But historically their output was usually just this cute thing that gave you a chuckle; they were seldomly as useful in a general sense like LLMs currently are. I think that the increase you mention in compute power and data is itself a huge step forward.
But also transformers have been super important. Transformer-based LLMs are orders of magnitude more powerful, smarter, trained on more data, etc. than previous types of models because of how they can scale. The attention mechanism also allows them to pay attention to way more of the input, not just the few preceding tokens.
If you want something useful, then we're getting closer.
AGI is something specific, as a requisite, it must understand what is being asked, and what we have now is a puppet show that makes us humans think that the machine is thinking, similar to Markov chains.
There is absolutely some utility in this- but it's about as close to AGI as the horse-cart is to commercial aircraft.
Some AI hype people are really uncomfortable with that fact, I'm sorry, but that reality will hit you sooner rather than later.
It does not mean what we have is perfect, cannot be improved in the short term, or that it has no practical applications already.
EDIT: downvoting me wont change this, go study the field of academic AI properly please
It seems clear to me that, if we could programmatically sample from a satisfactory conditional probability distribution, that this would be sufficient for it to, for all practical purposes, behave as if it “understands”, and moreover for it to count as AGI. (For it to do so at a fast enough rate would make it both AGI and practically relevant.)
So, the question as I see it, is whether the developments with ANNs trained as they have been, is progress towards producing something that can sample from a conditional probability distribution in a way that would be satisfactory for AGI.
I don’t see much reason to conclude that they are not?
I suppose your claim is that the conditional probability distributions are not getting closer to being such that they are practically as if they exhibit understanding?
I guess this might be true…
It does seem like some things would be better served by having variables with a fixed identity but a changing value, rather than just producing more variables? I guess that’s kind of like the “pure functional programming vs not-that” distinction, and of course as pure functional programming shows, one can still compute whatever one wants while only using immutable values, but one still usually uses something that is as if a value is changing.
And of course, for transformer models, tasks that take more than O(N^2) or whatever (… maybe O(N^3) because on N tokens, each is processed in ways depending on each pair of the results of processing previous ones?) can’t be done in producing a single output token, so that’s a limitation there..
I suppose that the thing that is supposed to make transformers faster to train, by making it so that the predictions for each of the tokens in a sequence can be done in parallel, kinda only makes sense if you have a ground truth sequence of tokens… though there is the RLHF (and similar) where the fine-tuning is done based on estimation of a score on the final output… which I suppose possibly neither is great at getting behavior sufficiently similar to reasoning?
(Note: when I say “satisfactory probability distribution” I don’t mean to imply that we have a nice specification of a conditional probability distribution which we merely need to produce a method that can sample from it. But there should exist (in the abstract (non-constructive) mathematical sense) probability distributions which would be satisfactory.)
In order for something to qualify as AGI, answering in a seemingly intelligent way is not enough. An AGI must be able to do the following things, which a competent human would do: given the task to accomplish something that nobody has done before, conceive a detailed plan how to achieve that, step by step. Then, after doing the first steps and discovering that they were much more difficult or much easier than expected, adjust the plan based on the accumulated experience, in order to increase the probability of reaching the target successfully.
Or else, one may realize that it is possible to reformulate the goal, replacing it with a related goal, which does not change much the usefulness of reaching the goal, but which can be reached by a modified plan with much better chances of success. Or else, recognize that at this time it will be impossible to reach the initial goal, but there is another simpler to reach goal that it is still desirable, even if it does not provide the full benefits of the initial goal. Then, establish a new plan of action, to reach the modified goal.
For now this kind of activity is completely outside the abilities of any AI. Despite the impressive progress demonstrated by LLMs, nothing done by them has brought a computer any closer of having intelligence in the sense described above.
It is true however, that there are a lot of human managers who would be equally clueless with an LLM, on how to perform such activities.
Use HN comments for comic relief. When you see one worth taking seriously, you'll know it. Otherwise we're just giving oxygen to contrarians.
Just to clarify, I'm definitely not saying "neuron simulation" is required in any way. I'm just asking, how can we be very close to "solving" a significant part of the most complex brains, yet miles away from solving the simplest brains?
You should be able to answer that question (or a steelmanned version of it), not just ridicule strawmen.
The immediate threat is that humans will use the leverage that LLMs give to replace and to influence other humans, in other words to gain power over us.
Whether this is AGI or not is beside the point.
As for the influencing part, what specific actions to gain power over us can be achieved now with LLMs, that could not be achieved before using a few tens of thousands of paid humans?
It's being able to do it without having to employ the tens of thousands of humans that makes it different. With the LLM you are able to react much faster and pay fewer people more money.
Empirically, it does not seem necessary to understand one version of a thing to produce a superior version. Why should the remaining unsolved cognitive tasks break that pattern?
Is this.....a reference to Greg Egan's Permutation City? Because that was the exact argument some characters used against real AGI(which let's assume they had, it's a Sci-Fi book). Basically it went along the lines of "even though we can simulate digestion at molecular level, nothing is actually digested. Why should simulating neuron activity create actual thoughts and intelligence?"
I do suspect however that there's something to the biological experience of "being the life support system" for the brain, that significantly affects the training process. It might be challenging to simulate that.
> In my view comparing Ai’s cognitive, creative or intellectual powers to those of the human brain is not especially helpful. Think of the car. Humans can’t run as fast as horses. But we can build machines that far outpace them. We do not achieve this by imitation. We don’t engineer mechanical legs and hooves of the kind that took evolution 34 million years of tinkering and modification from eohippus to the present day. We go a completely different way and we come up with something that doesn’t at all exist in nature: the wheel. And instead of a mechanical heart and mechanical muscles Karl Benz offers us the internal combustion engine and crankshaft. Ditto with flying, and travelling across or under the waves. The commonly held idea that the best engineering mimics nature is largely misguided. Yes, we look sometimes look to the natural world for inspiration but in the big things, structurally, we go our own way. And as a result we can fly higher and faster than birds, move over land quicker than a cheetah, swim over and under the water faster and further than a salmon or a whale and so on.
The biggest danger I see is a widespread AI with a set of badly defined goals, not a particularly smart and evil one.
> "there is absolutely zero evidence indicating that we are any closer to AGI than what James Watt was to realizing nuclear fusion"
James Watt lived before Rutherford split the atom, he didn't know they could be split or fused, he was not trying for nuclear fusion. We do know that information exists and can be processed. Still, James Watt was closer to large scale controlled release of energy than humans before the control of fire.
We know that human level intelligence is possible, in a way that Watt didn't know fusion was possible. We have looked for other mechanisms hiding in the brain - Penrose and Hameroff's ideas of quantum tubules for one - and rejected them. We've pretty closely bounded the amount of energy a brain uses, what it can sense, how it can input and output, what it's made of, and through what mechanisms of electric pulses and neurotransmitters it seems likely to operate. We've identified brain regions which cause predictable effects when damaged (on language, memory, executive decision making). We've dissected it, photographed it, x-rayed it, imaged it in layers, imaged it while active.
We've represented facts as data, written code which uses genetic algorithms to solve problems, written code which writes code to solve problems, written code which crunches data in large quantities to pull out higher dimensional patterns, code which appears to identify concepts from that, code which does specialised tasks such as face recognition, language recognition, language generation, person recognition.
It's pretty indefensible to say "absolutely zero evidence indicating we are any closer than people of 10k BC, people of 0 AD, people of 1800, or people of 1900". You're basically doing the God of the gaps argument, as we map the brain physically in more and more detail, and encroach from the data an information processing side, the gap between them where intelligence could be hiding appears to be shrinking. But with no evidence or support, you suggest something else is hiding there which makes the gap a Jaunt so large that encroachment from either side is immeasurably tiny. And you hide it behind the "G" for general.
Twenty years ago you could talk to a nonsense chatbot. Ten-ish years ago you could talk to a somewhat coherent GPT. Five years ago you could talk to a surprisingly coherent ChatGPT. Today people collaborate with Claude, ChatGPT, Gemini, ask them to explain concepts, provide examples and get them to follow on, offload work to them, learn from them, have them proof read documents, have them rewrite in different styles.
I think Hofstadter said he wasn't interested in computers because they don't have concepts. Here's an article in the ACM last year saying similar: ""Neural networks can look at a picture and tell whether it contains a dog, a cat, or a car, but they do not have a rich understanding of any of those categories. Take the concept of a bridge. Humans can extend the notion of a bridge to abstract levels. We can talk about a bridge between people or bridging the gender gap. We can instantly understand what these expressions mean because we have a rich mental model of what a bridge can be."" - https://cacm.acm.org/news/artificial-intelligence-still-cant...
Asking ChatGPT:
Me: "Hi, a bridge is a big metal thing over a road or river;
what does it mean to "bridge the gender gap"?
Can you explain briefly, please?"
ChatGPT: "To "bridge the gender gap" means to reduce the differences
or inequalities between men and women, particularly in areas
like pay, opportunities, education, or representation.
It's about creating a more equal and fair situation where both
genders have the same access and opportunities. Just like a
physical bridge connects two places, "bridging the gender gap"
connects the unequal sides, making them more balanced."
Me: "Interesting; can you suggest other non-physical things
which could be "bridged" please?"
ChatGPT: "Certainly! Here are some examples of non-physical things that can be "bridged":
Cultural gap – Reducing misunderstandings or differences between people from different cultures.
Communication gap – Improving understanding between people who are having difficulty communicating.
Generation gap – Narrowing the differences in attitudes and perspectives between older and younger generations.
Wealth gap – Reducing economic inequality between rich and poor.
Knowledge gap – Closing the gap in access to information or education between different groups.
Digital divide – Bridging the gap between those who have access to technology and the internet and those who don’t.
Trust gap – Rebuilding or enhancing trust between individuals or groups, like companies and their customers.
In each case, "bridging" is about connecting or bringing closer
two sides that are separated by some form of difference or inequality.
Seems to have some grasp of the general concept of "bridging" more than Eliza or Eugene Goostman, doesn't it? "I have tried some letter-string problems in GPT-3, and in some cases it could solve them. It learned, for example, the concept of successorship. Not perfect, not robust, but I found it still surprising that it can do this. Therefore, I don't agree that these systems are only 'stochastic parrots', as some scientists have called them. I have seen evidence of GPT building simple internal models of situations."" says the expert in the article.I'm not saying this is an AGI or will lead directly to AGI, but stubbornly refusing to accept this obvious progress as any progress towards machine intelligence at all, calling it "absolutely zero" evidence of progress seems wilfully blinkered.
Do you genuinely put us absolutely no closer, not a single step closer, to AGI than the Mechanical Turk or the people of 50k BC?
It’s still not much more than Markov chains, just with some clever anti-nonsense filtering and importance-weighting. There’s no “understanding”, nor anything particularly close to it.
It’s impressive we’ve accomplished so much with something that is so thoroughly, entirely stupid, in fact. They are useful tools, for sure.
what specifically would you or I do different, apart from having less training data?
> "There’s no “understanding”, nor anything particularly close to it"
any evidence for this claim? It explained, it responded in context to a followup question, it gave other relevant examples, by what measure does it "not understand" but I "do understand"?
I dunno, do they do more than that? Seems like it to me.
Does a Chinese Room “understand”? I say no, but hey, maybe it does.
If I laboriously do the math by hand, taking care never to actually know the informational content of any of the input or output myself, does my scratch paper understand your questions? If the output’s just as good as ChatGPT? Where’s the part that understands?
Yet somehow, from the connections of all of those cells, and neurotransmitters, there's consciousness and something there that does understand (and think and reason and love). If, instead of LLM architecture, on more powerful computers than we have now, we simulated all of those neurons and their connections, would we have a computer that understands? If we then did those computations on scratch paper, where would the piece that understands be on that piece of paper?
The sum of a thing's parts can be greater than the individual parts. Whether or not ChatGPT understands is a whole big question, but we'll have no more luck dissecting LLMs to find out if it does than if we dissected a human brain.
The calculations for operating an LLM definitely can be reduced to math. No reduction needed, in fact—they are math.
This isn’t an argument (to my mind, anyway) against even the possibility of machine whatever-you-like (consciousness, understanding, whatever) but against the idea of equivalence because we could simulate either one—in fact, we can’t simulate one. The other, essentially is already simulation, no further steps needed.
What you’re getting at (if I may attempt to present your argument) is that we could reduce either to its components and make it look ridiculous that it might be doing anything particularly advanced.
However, in fact we definitely can reproduce exactly what one of them does with a bunch of thick books of lookup tables and some formulas that we could mechanically follow by hand, and it might even be possible to do so in practice, not just hypothetically (at significant, but not impossible, expense) while we do not know we can do that for a human brain, short of just using exactly the brain that we want to “simulate”.
It isn't certain that it can be, but can you give any plausible reason why the Universe might allow understanding to (meat + electric patterns) and deny it to (silicon + electric patterns) ?
When I said "I'm not saying this is an AGI" and you reply with "I dunno human brains do more than ChatGPT" it feels like you haven't understood the discussion - that part was never contested.
> "some clever anti-nonsense filtering"
Eliezer Yudkowsky wrote 'The opposite of intelligence isn't stupidity'. On an A/B test, stupid is guessing randomly and that approach scores 50%. Scoring 0% takes as much intelligence as 100% because it requires knowing the right answer to be able to avoid it every time.
Being able to identify nonsense is sense. At the risk of being tautological, "clever" is clever.
If the neurons in your brain are the scratch paper in the Chinese room, each one isn't aware of the content of the light waves or the finger muscle signals, and you conclude the Chinese room doesn't understand, shouldn't you conclude that your brain doesn't understand? If your brain does understand shouldn't you conclude the Chinese room would understand?
I claim that ChatGPT being able to explain bridging and give further examples is behaviour which demonstrates more understanding than a rock, than a calculator, a wordlist, a spellchecker, a plain Markov chain has.
You say there's "no understanding or anything close to it" - how would ChatGPT's response look different if it did understand the concept of bridging?
If you cannot suggest any way its output would look different to how it looks now and instead have to resort to changing the subject, shouldn't you retract that claim?
> "Does a Chinese Room “understand”? I say no"
Then you must say a human doesn't understand. For what else is there in a human brain except a finite amount of learned behavioural rules for signal inputs and outputs? Learned over a billion years of evolution in the structure, and filled in by a lifetime of nurture.
For one thing, I expect we’d not see so many cases of them chasing (if you will) the prompt and request into silliness. The code attempts to satisfy prompts in a transparently mechanical fashion, which is part of why they so gleefully (if you will) mislead. There’s no understanding. You can ask them to correct and they might, but they can also be induced to correct the already-correct, so that means nothing. To the extent we fix that, it’s not by adding any factor that might represent understanding, it’s further prompting that amounts to “follow these patterns slightly differently”. The fix isn’t, so far, “teaching” them to understand. Maybe we’ll get there! But we don’t appear to be anywhere near that yet.
> Then you must say a human doesn't understand. For what else is there in a human brain except a finite amount of learned behavioural rules for signal inputs and outputs? Learned over a billion years of evolution in the structure, and filled in by a lifetime of nurture.
The thing about the Chinese Room is that we comprehend the entire process, and there’s no room for some unknown factor affecting the output—or for a known factor that might be processing something like what we mean by understanding (let alone consciousness, say).
Every single part of what an LLM does can be replicated with big books of lookup tables, dice, and a list of rules. There’s nowhere for anything to do the understanding to exist. It’s not that we have to be confused by part of it for that to be there—I’m not saying mystery is a necessary component—just that this process doesn’t have a place for that to be.
In the one example I gave you saw one output and declared it "not understanding the concept of bridging". I'm asking specifically that output, how would it look different if ChatGPT had some understanding of the concept of bridging? You're back to arguing "it's not human level!" which was not my claim. My claim is that it's above zero level. In another comment I asked it to use the concept of bridging in new ways, and it provided sentences which have no hits on Google but are plausibly the kind of thing I might see in a book from a human author.
> "There’s no understanding"
Say to your pet "I like it when you do human-like things such as standing on two feet. Come up with more human-style things for more treats" and it won't. You can ask ChatGPT to come up with more uses of the bridging concept, and it does. That is demonstrating understanding at higher than rock level and higher than rat level, and you can't reject that evidence just by repeatedly saying "there's no understanding there's no understanding there's no understanding".
> "they can also be induced to correct the already-correct, so that means nothing."
So can I; if my boss tells me there is an error and I need to correct it, I might correct a non-error to please them. Knowingly ("I'll change this part from correct to wrong if that pleases them") or unknowingly ("if they tell me there is an error there must be one, I'll take a guess that this bit is wrong and put something else here"). Does that show I have no understanding?
> "Every single part of what an LLM does can be replicated with big books of lookup tables, dice, and a list of rules. There’s nowhere for anything to do the understanding to exist."
You're doing the God of the Gaps argument with the human brain - an LED screen is RGB pixels, there's nowhere for a picture of a cat to exist separately from bright and dark pixels. A book is printed characters, there's nowhere for a story to exist separate from blobs of ink on paper. A brain is meat grown from a foetus, uses ~20 Watts of energy, if the blood supply is cutoff then it dies, if it gets too hot or cold then it dies, there are many areas which can be damaged and harm something like leg movement but there is no single area which can be damaged which stops 'understanding' but leaves everything else unchanged, there are no examples of people being decapitated, having no brain, having brain death, and still having 'understanding' provided by whatever other thing you are implying exists and does understanding.
There's nowhere for anything to do the understanding to exist, unless there is a) new physics which aligns perfectly with every observation we have about the brain but also augments it an adds some magical 'understanding' thing which can't be done or simulated in software. b) something non-physical such as a soul which is tied closely to the meat and powered by the food and blood and can't be tied to silicon because reasons. c) ??? As far as I can see this isn't reasoning from anything more convincing than you not wanting to accept the Occam's Razor simpler explanation that a purely physical information processing system can understand.
(Or that humans don't understand and it's all some weird illusion; the picture of the cat is not in the LED screen, it is in the eye of the beholder. The understanding isn't in your behaviour, it's in the beholder's interpretation, I believe you understand because you demonstrate the behaviours of understanding. We are seeing intelligence in others where there isn't any. And that view turned on ourselves is our own perception of our own understanding - we see ourselves identifying patterns, extrapolating patterns, continuing coherent sentences, and conclude that we must have 'understanding' as a thing separate from those behaviours).
> "The thing about the Chinese Room is that we comprehend the entire process, and there’s no room for some unknown factor affecting the output"
We don't comprehend the entire Chinese Room; the instructions that Searle is following are a massive handwave. Does following the instructions require Searle to make human judgements on where to branch? Then it's offloading understanding onto his human brain. Does it not require that but it still outputs coherent responses? Then the instructions must encode intelligence in them in some way - if intelligent behaviour doesn't demonstrate intelligence we're in non-scientific nonsense land.
Peter Cochrane wrote about 'dying by installments' of a human turned into a cyborg replaced bit by bit, Ship of Theseus style. We can do similar and make up a Cochrane's Chinese Brain - instead of a neuron firing and affecting the connected neurons, it raises an alert and Searle walks over and writes down the firing pattern on a scratch pad, walks to all the other relevant neurons, and taps in the firing pattern on an input device, without understanding the information content of the firing pattern. Does the brain keep responding coherent Chinese but no longer understand Chinese?
Let’s try this:
We could apply an LLM to made-up language and corpus that does not actually carry meaning and it would do exactly what it does with real languages.
“Well maybe you accidentally encoded meaning in it. We could always, say, cryptoanalyze even an alien language and maybe be able to come up with some good guesses at meaning”
Maybe we could. But now imagine also you have no “knowledge” whatsoever except the trained patterns from that language. Like, no understanding of how to do cryptoanalysis, or linguistics, or what a planet is. Or an alien. All you’re doing is guessing at patterns, based on symbols that you aren’t even attempting to understand and have no basis for understanding anyway. That’s an LLM.
I think people are assigning way too much power to language sans… all the rest of what you need to derive meaning from it. None of what’s going into or coming out of an LLM needs to carry any meaning for it to do exactly the same thing it does with languages that do.
To the extent that an LLM has a perspective (this is purely figurative) all languages are gibberish alien languages, while also being all that it “knows”.
> We don't comprehend the entire Chinese Room; the instructions that Searle is following are a massive handwave. Does following the instructions require Searle to make human judgements on where to branch? Then it's offloading understanding onto his human brain. Does it not require that but it still outputs coherent responses? Then the instructions must encode intelligence in them in some way - if intelligent behaviour doesn't demonstrate intelligence we're in non-scientific nonsense land.
I remain stubbornly unconvinced that simulating a real process (by hand or otherwise) is the same thing as it actually happening with real matter and energy, even setting aside that the most efficient way to achieve it is to… not simulate it, and use real matter to actually do the things.
It’s why I find the xkcd “what if a guy with infinite time and an infinite beach and infinite rocks moved the rocks around in a way that he had decided simulated a universe?” thing interesting as an example but also trivial to solve: all that happens is he moved some rocks around. The meaning was all his, it doesn’t do anything.
You opened by saying you aren't doing God of the Gaps, but here you are doing it. Brains move chemicals and electrical signals around. That doesn't do anything, apparently. Matter doesn't do understanding. Energy doesn't do understanding. Mathematical calculations don't do understanding. Neural networks don't do understanding. See how Understanding is retreating into the gaps? Brains must have something else, somewhere else, which does understanding? But what, and where? It's a position that becomes less tenable every decade as brains get mapped in finer detail leaving smaller gaps, and non-brains get more and better Human-like abilities
> "there’s both nothing we know of doing understanding .. it’s not doing understanding."
It is. The math and the training and the inference is the thing doing understanding. Identifying patterns and being able to apply them is part of what understanding is, and that's what it's doing. [Not human level understanding].
> "We could apply an LLM to made-up language and corpus that does not actually carry meaning and it would do exactly what it does with real languages."
We do that with language too; the bouba/kiki effect[1] is humans finding meaning in words where there isn't any. We look at the Moon and see a face in it: Pareidolia[2] is 'the tendency for perception to impose a meaningful interpretation on a nebulous stimulus so that one detects an object, pattern, or meaning where there is none'.
We are only able to see faces in things because we have some understanding of what it means for something to 'look like a human face'. "We see a face where there isn't one" is no evidence that we don't understand faces and so "an LLM would find patterns in gibberish" is no evidence that LLMs don't understand anything.
> "All you’re doing is guessing at patterns, based on symbols that you aren’t even attempting to understand and have no basis for understanding anyway. That’s an LLM."
Trying to build patterns is "what attempting to understand" is! You're staring right at the thing happening, and declaring that it isn't happening. "AI is search" said Peter Norvig. The Hutter Prize[3] says "Being able to compress well is closely related to intelligence as explained below. While intelligence is a slippery concept, file sizes are hard numbers. Wikipedia is an extensive snapshot of Human Knowledge. If you can compress the first 1GB of Wikipedia better than your predecessors, your (de)compressor likely has to be smart(er). The intention of this prize is to encourage development of intelligent compressors/programs as a path to AGI". Compression is about searching for patterns.
Understanding is either magic, or it functions in some way. Why not this way?
> "all languages are gibberish alien languages, while also being all that it “knows”."
If we took some writing in a Human language that you don't speak, you can do as much "predict the next word" as you want, take as much time as you need, and put together as an output. The input is asking for a reply in formal Swahili which explains yoga in the style of Tolkein with Tourette's, but you don't know that. The chance of you being able to hit a valid reply out of all possible replies by guessing is absolutely zilch. But you couldn't do it by " predicting the next word" either, how would you predict that the reply should be in Turkish if you can't understand the input? How would you do formal Turkish without understanding the way people use Turkish? Conversely if you could hit on a good and appropriate reply, it would be because your studying to "predicting the next word" had given you some understanding of the input language and Swahili and yoga and Tolkein's style and how Tourette's changes things.
> "I remain stubbornly unconvinced that simulating a real process (by hand or otherwise) is the same thing as it actually happening with real matter and energy"
Computers are real matter and energy. When someone has a cochlear implant, do you think they aren't really hearing because a microphone turning movement into modulated electricity is fake matter and fake energy, and an eardrum and bones doing it is real matter and real energy? Yes it's true that you can't get on a simulation of a plane and fly to New York, but if you see the output of an arithmetic calculation there's no way to tell if it was done with a redstone computer in Minecraft or with Python or with brain matter. (Is it possible for arithmetic to be not-simulated?).
[1] https://en.wikipedia.org/wiki/Bouba/kiki_effect
No! There’s a difference between a thing happening, and symbols we decided mean something bearing manipulated. The assigned meaning isn’t real in the way an actual process is. A flip-book of a person jumping rope isn’t a person jumping rope.
Create things that are not in the training data. I could coin a new phrase that uses "bridge", I doubt an LLM can or would have any motivation to. Because LLMs understand data and patterns, but they don't understand the meaning behind them. They can't be imaginative.
For example, an LLM can recite Shakespeare and can even write new works in Shakespeare's style. But it cannot be a new Shakespeare, the is one level above it. An image generator might be able to make a painting is Picasso's style, but it cannot be a new Picasso - that is, create a style never before seen and explain why they did it and what emotions it conveys.
Me: Hi, can you use the concept "bridge" in a
never-seen-before sentence, where it still makes
sense? Please make the sentence very unusual,
but still coherent. Perhaps a couple of examples?
ChatGPT: Sure! Here are a couple of unusual but coherent
sentences using the concept of "bridge":
The stars whispered their ancient stories across the
sky, building a bridge of forgotten time between my
dreams and the universe’s memories.
Each thought she left unspoken laid a ghostly bridge
between us, invisible but weighty, made of all the
words that could have been.
Both sentences play with abstract ideas, pushing the
concept of a bridge beyond its typical use but
maintaining coherence through metaphor and vivid
imagery.
----It's impossible for me to know everything in the training set but there are no Google results for "building a bridge of forgotten time" or "laid a ghostly bridge between us". If you told me those and said you made them up and didn't copy them from somewhere, I'd have to either find the source you copied or believe you.
Its patterns encoded two items and “between” being likely to occur, anyway.
---
I asked ChatGPT to help you better understand how it works:
There are a few common misconceptions in that comment regarding how large language models (LLMs) like GPT-4 actually work, so let's clarify those:
Markov Chain Comparison:
LLMs are not based on Markov chains, though they might seem similar at a high level due to their ability to predict the next word in a sequence. Markov chains rely on simple probabilistic transitions between states, often based on a very limited "memory" of previous states (e.g., the previous word or two). LLMs, on the other hand, use a transformer architecture, which allows them to consider long-term dependencies and relationships in text. This means they can account for the context of many preceding words, sentences, or even paragraphs when generating responses. No "Understanding":
While it’s true that LLMs do not have consciousness, self-awareness, or human-like understanding, the term “understanding” can be misleading. They operate by modeling patterns in language, but in a highly sophisticated way. LLMs capture a deep representation of the relationships between words, sentences, and broader concepts through billions of parameters, giving them a kind of statistical "understanding" of language. This enables them to generate coherent and contextually appropriate responses, even if it’s not the same as human comprehension. Importance Weighting and Search:
LLMs do not search through predefined sets of phrases or apply “importance-weighting” to words in the way described. They generate text dynamically by using the probabilities derived from the training data they’ve seen. The model calculates probabilities for each possible next word in the sequence, taking into account the entire context (not just key terms), and selects the next word based on these probabilities. This process is not about tagging words as important but about predicting the next most likely word or phrase given the context. Not Just "Anti-Nonsense Filtering":
The quality of LLM output doesn’t arise from filtering out nonsense but from the underlying model’s ability to capture the complexity of human language through its learned representations. While there's a certain degree of training that discourages incoherent outputs, the coherent responses you see are mostly due to the model's training on vast, high-quality datasets. "Thoroughly Stupid":
It's more accurate to say that LLMs are highly specialized in a particular domain: the patterns of human language. They excel at generating contextually relevant responses based on their training data. While they lack human-style cognition, calling them "stupid" overlooks the complexity of what they achieve within their domain. In summary, LLMs use advanced neural networks to predict and generate language, capturing sophisticated patterns across large datasets. They don't "understand" in a human sense, but their ability to model language goes far beyond simple mechanisms like Markov chains or weighted searches.
(The “All You Need is Attention” paper is fairly readable, all things considered, and peels away a lot of the apparent magic)
No, don't get me wrong, I absolutely acknowledge that we have made progress and can produce very useful things that are rightfully called machine intelligence! And probably there are things we are figuring out now, that will be relevant and useful even if we someday figure out AGI.
I specifically chose Watt as an example because he also produced a very useful thing that improved the world. And many concepts from that time are still used today, even if we don't have many steam engines anymore.
That he didn't have the concept of fusion is beside the point - we have many examples of cases where we have the concept, but will not be able to achieve it in thousands of years (like Level 2 on the Kardashev scale). And vice versa, where we go from discovering concepts to real world impact in just a few years (like GPTs).