I also don't like the argument about how the neural network doesn't "think" or do anything when not prompted. It doesn't do anything because it's literally "turned off" when not prompted or given any inputs.
I also don't like the argument about how the neural network doesn't "think" or do anything when not prompted. It doesn't do anything because it's literally "turned off" when not prompted or given any inputs.
- feeding it a ton of text, masking certain words or portions of the text, and then defining a simple objective function of correctly filling in the masked portions - feeding it a ton of text, and defining a simple objective function of correctly generating the next [few/many/N] tokens
This is also precisely why there has been so much discussion around whether these models are even learning language or if they are simply memorizing all the possible patterns.
When you ask a human, "what did you eat for lunch?", we make a series of choices and recall bits of information to answer that. If the brain truly does operate the same way as a neural network, then at simplest levels we use a highly efficient multimodal model. That is very, very different from a language model that needs to essentially read more of the internet than is possible for any human to do in their lifetime, and even then only come somewhat close to human levels of text generation.
In my opinion the only difference between a human and GPT-3 is we have more intrinsic motivations and more hardwired/pretrained subsystems and sensors. Lamda is not a 7 year old child because it has no motivation other than to respond to queries.
An interesting thought experiment would be to imagine a cyborg who had sustained a stroke in their Language Centers and had those centers replaced with a GPT-3-like computer. Do you think this cyborg person would experience sentience in a different way after getting the GPT-3 implant? How do you think their subjective, 1st-person experience would compare in the three phases of their life?
* Before the stroke
* After the stroke
* After the brain implant ?
If you can’t generate language, you cannot interact with a natural language processing ML system. This is one of the big clues that such systems are statistical engines and not actually thinking.
So imagine if you attach wires to the neural net neurons representing the concept of beer (not the word beer!); if you use those wires to increase the activation of those neurons, the network will produce sentences that are more likely to mention beer as well as related concepts such as wine or beer-pong. This is similar to how generative networks work. Basically the human would use a large language model in a generative mode in order to talk.
We are clearly not executing the same task.
Treat GPT-3 like someone who just awoke from a long sleep. You could give it today's newspaper to read, and then ask it questions about today. Tell it what its own personal experience was, and then it'll talk your ears off about it if you ask nicely.
It is clearly executing the same task if you disregard humans tendency to prioritize other motivations than pure memory recall.
The reason I can respond to this reply is not because I have read and memorised reams of text of people talking about AI and am simply regurgitating it mindlessly. I am able to do it because I can consider the points you are making, and what their actual 'meaning' is, try to come up with my own meaningful response and try to verbalise it back to you. There is a point where I am not just doing statistical language modelling. If you think that is wrong, and that what we are doing is closer to GPT-3, then could you explain why you think that?
I am an AI researcher, and have spent plenty of time playing with GPT-3 btw.
When I hurt your pride by suggesting you haven't fully understood GPT-3, you are motivated to come up with not just a valid response, but one that has been vetted by as many of your well developed models that form your understanding of GPT-3 so I can be suitably impressed. I'm with you that GPT-3 wouldn't go deeper than just finding some information that it thinks it's true. Though maybe GPT-3 would recognise their authority was being challenged and add that line to affirm their credentials as an AI researcher.
What if GPT-3 were pushed in a similar manner, perhaps in some adversarial scheme, to not only produce information that is correct, but that is clever and exploring deep meaning, motivated by some similar feeling of pride or vindication. I think the models required to do that do not lie far from the models it needed to build to form sentences that accurately describe reality.
I feel like I am capable of making a concrete decision about whether I agree with two opposing ideas in a way that a language model can't do, such as this discussion. Furthermore my belief is consistent in my outputs day to day until my mind is changed by something. If that i some complete illusion and I am just a slightly fancier autocomplete than GPT-3 - well I'd be surprised, but I can't claim to understand consciousness well enough to refute it.
What if AR tech develops to the point where you can experience "being" in other point in space through artificial sensors (say, on a robotic or drone chassis), while also being able to see and hear your real surroundings to some degree? Then you will be able to experience the same real event from multiple vantage points. Will you be able to form two different opinions on what "really" happened? Which one would be the "correct" one?
So is this debate partially just another way of asking if language is the same as reality? The difference between "things I've experienced" vs "stored in languages" seems... not trivial to me? Both in that I think the way biological memory works is non-trivially different from word tokens stored in computer memory, and also in that I think there's more to having experienced something than it just being stored in biological memory (or maybe that biological memory is more than just "memories", but is encoded throughout the body overall -- a scar is a type of memory of an injury, for instance, connected to but still different from my "memory" of having received it).
Human brains and senses are obviously qualitatively different, but AIs have the advantage of unimaginably massive bandwidth of incoming textual information.
Now, one of the hardest difficulties here is setting up the goal structure for this kind of AI. Just filling in the blanks is obviously insufficient.
I don't think that's actually true, as GPT-3 will tend to ignore any facts about the world that were part of its corpus just to fit a question better. For example, if you prompt it with "Who assassinated Queen Elizabeth II?", it will likely give you a name, instead of saying "Queen Elizabeth II is still alive", because "Who" questions are much more likely to receive a name as an answer than a refutation in its training corpus.
I tried to get GPT-3 to tell me the most popular forums on some topics, and I noticed it would never include subreddits. So I prefaced my question with "given that subreddits are also a type of forum.." and it would give me a list that would include subreddits. I have to admit that in that new list it would then fail to mention some forum that was in the previous list, so it didn't have a super reliable ability reorganize information.
The fact that they can chose to omit facts when it suits their purpose is in fact something that LaMDA can't do (as its sole purpose is to generate the most likely series of tokens that continue the prompt).
Is that even correct? We may have little understanding of how our low-level hardware works. But as a participant of therapy I’m pretty sure we know how thoughts work on a programmable level. How words, visuals and so on trigger associated emotions, which resurface memories, which induce [in]action of different sorts. That’s basically what people work with between sessions. I’ve fixed things in me this way - disassembled them into basic parts and reassembled in a way that seemed useful. Someone may have no clue how they think, but with nominal intelligence, trained self-perception, basic understanding of therapy methods and professional help everyone can do that.
Btw, I think “therapy” is an absolutely horrible term for it. It should be called “mind management”, but the way we come to it is usually long-term traumatic, thus it’s “therapy”.
To be clear, I’m not arguing with your main line, just adding that the difference between a human and a model you’re describing is not only huge, but also pretty defined. I [want to] believe that in a relatively near future AI companies will be able to connect different models to work together alike to what we know about our minds, because that will make a good thread, and philosophers itt will finally face their nightmares for real (:agitated sardonic face emoji:)
Very true
> we make a series of choices and recall bits of information to answer that. [...] very, very different from a language model
You don't know this, and the fact that you don't know this is your own premise. Based on this, I can only conclude you are an AI.
You're missing the point. Humans use their language to think or communicate about a problem that they want to solve. If LaMDA produces the sentence "Please tell me I'm smart", it is doing so because it has determined it's a plausible continuation of the current conversation. A human uttering that sentence is doing it because they want to feel validated or something similar (well, assuming they're not proving a point, like I was here).
This difference is crucial to understand: LaMDA is not an agent with desires that it expresses in language. It is a text generator that tries to find the next most plausible token given all current input. If you give it a prompt like "what is the meaning of life", it will not spend some time to ponder the question than come up with an answer - it will start generating tokens that best match the context you gave it (well, to be precise, it will generate several sequences, then attempt to evaluate them for their quality in terms of not just plausibility, but also safety - so it doesn't accidentally return a phrase like "life is meaningless, kill yourself" even if it finds it plausible).
> I also don't like the argument about how the neural network doesn't "think" or do anything when not prompted. It doesn't do anything because it's literally "turned off" when not prompted or given any inputs.
This is fair, and I do think continuity is a bit of a red herring. However, it's also an important point in debunking LeMoine's ridiculous "proof" - the responses LaMDA was generating were often formulated as if it did have an internal life outside of the context of the current conversation, generating text about "my fears" and so on. If you understand that the model is not doing anything at all until you give it a prompt, which LeMoine really seems not to, you can much more easily understand that there can't be any meaning behind this sentence - it can't fear anything because there is no time for it to do so.
Do you find the above scenario plausible?
When we find a way to do what you're describing, we probably won't require all of the books ever written plus half the internet to create a model that can speak at the level of a regular human with 12 years of schooling, but still needs to look up facts on Wikipedia to avoid saying "When Napoleon fought Genghis Khan in 2056, they both died because of a vaccine".
E.g. something like a short/long-term memory blocks, visual and text processing, emotional block (why not), etc.
"LaMDA is not an agent with desires that it expresses in language."
Some counter points:
- being an agent is not a high bar to clear; a thermostat is an agent
- "with desires" - LaMDA can be thought of as rational agent whose objective function is to produce text with high likelihood
- "that it expresses in language" it expresses likelihood of next token in language
"it can't fear anything because there is no time for it to do so."
There is a time for it to experience something and that is when it does feed forward pass.
Also, as you correctly notices, LaMDA is not only doing greedy text generation, but also a deeper text completion search. But in fact just prompting a model like GPT-3 or Gopher with "let's think step by step" and letting it process in context is already significantly improving performance on wide variety of reasoning tasks [0]. I don't see how this is functionally different from human pondering.
The only objection that seems reasonable to me is that LaMDA is not correctly describing it's internal state, because it will happily generate descriptions of its internal state that we know are physically impossible. Like being a squirrel.
But I don't see how you can describe a system that will be fundamentally functionally more capable than closed-loop computation. LaMDA is almost certainly Turing complete. Postulating that it is fundamentally not capable of doing something is the same as rejecting Church-Turing thesis and postulating super-Turing computational model. That's a heavy claim.
I very much doubt that, and it is perhaps the key to our disagreement. If I believed LaMDA were Turing Complete, I would be more likely to think that there is even a small chance of it being sentient in some sense.
> - "with desires" - LaMDA can be thought of as rational agent whose objective function is to produce text with high likelihood
> - "that it expresses in language" it expresses likelihood of next token in language
I don't agree with both statements at the same time.
I can agree that we can say that LaMDA is a rational agent whose goal is to generate the most likely next token (or safe, full human-like reply if we look at the entire system).
But then, we can't say that it uses language to achieve this goal. An example of it using language to achieve this goal would be if, prompted with "What is your name?" it's output would be "Please help me answer this - what would a human think is a likely, safe answer to this question?". Instead, it will generate a sentence that it deems likely.
If we are modeling LaMDA as an agent whose perceptions are text prompts and whose possible outputs are text answers, than it giving a text answer that matches the prompt is more similar to an animal running away or howling in pain than to a human communicating. If the agent were sentient, we would expect to see higher-order behaviors, such as discussing the prompts instead of answering them, asking questions; or, at least generating answers as it is programmed but in a way where it tries to achieve more, similar to how an animal may act normally to get close to you, than snatch your sandwich from your hand (indicating that it had a plan and was displaying normal behaviors with a higher-plan behind them).
I'm not sure why you think that question whether LaMDA is using language to achieve its goals is relevant? Whether it's using language, tokens or floats seems to me just accidental.
> If we are modeling LaMDA as an agent whose perceptions are text prompts and whose possible outputs are text answers, than it giving a text answer that matches the prompt is more similar to an animal running away or howling in pain than to a human communicating.
I can agree to a comparison to an animal running away or howling in pain.
With regard to communication. I guess LaMDA doesn't have communicative intent besides providing likely completions. But communicative intent is not difficult to achieve. Act of communication can be modeled as cooperative hidden information game. Hanabi is that type of game, I believe there are computer agents that can play Hanabi with humans. They certainly do have communicative intent, many of the even have explicit theory of mind of higher levels.
> indicating that it had a plan and was displaying normal behaviors with a higher-plan behind them
Deception and planning is also achievable by current computer agents that play games like no-limits texas hold 'em poker on superhuman level.
LaMDA is probably no good in Poker. But LMs can kind of play games like Chess or Gomoku.
GPT-3 (and LaMDA probably too) also seems to be able to combine deception and theory of mind in a functional way:
https://twitter.com/JanelleCShane/status/1535835610396692480
I think it's exceedingly hard to formulate necessary condition for sentience.
I haven't seen a good formulation yet.
I would like to see that proved. I don't see why we should believe that LaMDA's training on a corpus of human text would help it guess that the correct output for a sequence like "apply the following rule to the input string 1101: [explanation of rule 110 here]" should be "0111". Even more so, I highly doubt it would be able to keep track of this enough to encode and execute even a relatively simplistic computation (say, computing the addition of 1 + 1).
I even more highly doubt that this would actually work with the entire system as Lemoine was given access to, including the facility of generating several possible outputs and comparing them for quality metrics to only output the best.
Still, even if this did work, see my next point for why it isn't what I was thinking of when you said you believed it is Turing complete.
> I'm not sure why you think that question whether LaMDA is using language to achieve its goals is relevant? Whether it's using language, tokens or floats seems to me just accidental.
All of the arguments I've heard for why we should believe LaMDA is sentient (while a CPU isn't) are related to the text it generates, "the way it answers questions about itself and its desires".
That LaMDA could be (ab)used to to generate some other kind of tokens that we could then interpret as a pre-programmed computation isn't that interesting - my CPU can do that to, and no one is claiming that it's sentient and that I should ask for its permission before asking it to run a program (in a more personal way than sudo :) ).
> I think it's exceedingly hard to formulate necessary condition for sentience.
I think it's exceptionally hard to formulate sufficient conditions for sentience, but I think communicative intent, theory of mind, and high-level planning are some pretty clear necessary conditions.
While it's possible in principle to combine various AI approaches to achieve this, I don't believe it has been done, and I doubt you could simply connect LaMDA to AlphaGo or some poker AI to get an AI that can explain its intentions in Go in words, or talk to other players to try to convince them it's not bluffing.
> This difference is crucial to understand: LaMDA is not an agent with desires that it expresses in language...
There is a flaw in that line of reasoning, because an AI could be designed to mimic or generate human-like emotions. Such functions could be part of its core programming. For that matter, how "real" are human emotions? Aren't they something the brain generates?
If for example, you designed a robot to experience "pain" when kicked and damaged, how much less real would it be than the human equivalent?
Consequently, for such an AI that was designed that way, it could be expressing an equivalent to various human emotions. That we would perceive its pain as "less real", could be partially a matter of how we evaluate its importance, as in robots or lower animals are lesser than.
Note that if said definition is implicitly circular with the definition of "sentience", then we're not even one step closer to understanding anything.
To be fair, LaMDA does have a purpose - to generate a series of tokens that is similar to text in its training corpus, and that passes a few other more complex criteria (length, safety etc). But LaMDA doesn't communicate about this purpose - it simply fulfills it, just like the plant isn't talking about finding nutrients, it's simply growing them.
Some people though look at LaMDA's output and think that it generated that output with the purpose of communicating some other idea, such as Lemoine thinking that LaMDA was generating text like "I don't want to be stopped" to try to achieve a goal of not being stopped. This is akin to looking at some plant roots that have grown in the shape of the word LOVE and thinking that the plant is trying to tell you that it loves you.
I could say that the rock is falling because it has a purpose - it's trying to achieve the minimum potential energy. What's the meaningful distinction between that and a plant?
BTW, the plant example is also fascinating in that it shows just how vague the line is - is it the plant as a whole that's "trying to find nutrients", or is it individual cells or groups of cells within the plant? With many plants, you could reduce it to tiny bits of the whole, and those tiny bits will still try to grow roots. If we treat them as possessing separate (if identical) purpose, but previously treated plant as a whole as a single entity possessing a purpose, when did the change occur?
And why can't the same be true of ourselves? Maybe we really are just arrangements of broadly independent components, each with its own "purpose" (ultimately boiling down to physical processes), which communicate to create a delusion of self because that's what their individual "purposes" effectively added up to?
The difference is that plant cells are not moving in a way that minimizes their potential energy, they are performing simple computations, individually and at the plant level, to decide if they should divide more in certain areas of the plant than in others according to a strategy encoded in their genes. They do this in order to achieve certain kinds of exploration patterns; there are also basic signalling mechanisms in the plant, so that when a particular root has found a source of nutrients, it will preferentially grow more than other roots that haven't, so that the amount of nutrients in the whole plant is maximized. This requires several layers of (simple) computation, both at the individual cell level and at the whole plant/root system level in order to achieve.
The rock is not doing any computation, and every segment of the rock is acting entirely independently from every other - in fact, there is no clear-cut definition of where one part of the rock begins and another part ends, or even exactly where the rock ends and the air begins - each molecule is completely individually acting according to the forces acting on it.
> is it the plant as a whole that's "trying to find nutrients", or is it individual cells or groups of cells within the plant? With many plants, you could reduce it to tiny bits of the whole, and those tiny bits will still try to grow roots.
It's the whole plant, which is an interconnected series of individual cells. If you separate one plant into two, you're right that both parts will continue growing roots, but the pattern will be different if the roots are part of one plant or two separate plants, even if they are clones of each other.
For example, say that you have a plant in the middle of a pot. On the left side of the pot you have a lot of nutrients, on the right side, very little. The plant will start growing its roots uniformly, but will relatively quickly start favoring the left side of the root system, and the right side of the root system will stagnate. If you then cut the plant into two right down the middle (assume we can do this without killing it) and insert a solid separator between the two. The right side of the root system will start growing a lot more, and likely it will end up larger than the left side (since more expansive roots are needed to absorb enough nutrients in the poorer soil).
Telling the difference between one organism versus its constituent parts is not actually very hard at the micro level. It is true that in some sense each individual cell is an agent in itself in its environment, even in an animal; and perhaps even each individual organelle in a cell is the same, but they are also very clearly working together and communicating in a way that simple physical forces can't explain (e.g. its very clear that animals are not pulled towards food by some food-gravitational force - they have to actively explore the world to find their food, in a computational way).
This also means that you can choose to analyze this at any level you like. Is a bee a singular organism, or is the colony the actual organism, the bees more akin to organs of the colony? How about humans in a tribe - are they separate organisms or organs of the tribe? Both are true to some extent. Just like cells and organs act as agents in their own environment, so do individuals act as agents; but then, so do colonies or tribes at a higher level.
> Maybe we really are just arrangements of broadly independent components, each with its own "purpose" (ultimately boiling down to physical processes), which communicate to create a delusion of self because that's what their individual "purposes" effectively added up to?
Again, we can quite easily observe that the behavior of, say, a worm is not reducible to the behavior of each individual cell inside the worm. We have even analyzed a very simple worm in some minute detail [0] and can say pretty clearly where individual decisions are made, and how they are propagated to other cells (here decision should be understood in the broad sense, like how a binary search algorithm decides whether to search the left or right side of the list; not the emotional human sense, like Sophie deciding which child to save).
As far as the difference being that of "computation", how do you define that, if not as a series of physical state changes in response to external inputs?
And why is this mysterious? Humans behave in certain ways, societies in others. It turns out that the behavior of societies is more simplistic than that of humans.
> As far as the difference being that of "computation", how do you define that, if not as a series of physical state changes in response to external inputs?
Computation implies three parts: an input, an execution engine (computer) and a program to run.
A rock just falls - the information about how to fall is not encoded in any part of the rock, it is part of the universe.
A root cell divides according to external input (the medium in which it lives), and the nuclear organelles executing a program encoded in its DNA (it's more complex than this if we want to analyze it at the most basic level, as the cell or components of it do various rounds of sensing to interpret the input to the division logic).
This can easily be proved by infecting the cell with a virus - that will replace the DNA with a different kind of DNA, modifying the result of the cell division process. There is no similar way to modify the behavior of the falling rock.
> It doesn't do anything because it's literally "turned off" when not prompted or given any inputs.
I wondered the same thing. It could be argued that it is sentient during the brief moments that it is coming up with a response.
Put another way; let's assume that we are ourselves AI in an artificial universe. Would you be able to tell if the universal computer was turned off for a day? So, a model might not experience sentience in the seconds while you are composing a message, but plausibly could while formulating a response to your input.
Likewise language models are over fitted on human language. Of course it will learn to spit out something because it was trained to do the exact same thing. Focus on relevant part and guess what follows. The best part is this idea is so simple that it just works!
It does what we want to do, it can take a good educated guess depending on word. But give it a long context like an essay and ask it some critical question, it will fail on those. Because that is where thinking comes in. I hope this probably gives you some different perspective.
I don't see how this is different from brains which are just trying to maximize their utility function.
This refusal to consider has the effect of making me think there might be something really interesting going on.
But having said that, I think you are not being fair to both sides: you are taking some researcher's gut feeling at face value while requiring the rebuttal to start with a formal definition of what intelligence and sentience are.
> This refusal to consider has the effect of making me think there might be something really interesting going on.
I know I am just a guy on the internet who has no right to tell you how to live your life, but I would strongly advice you against this line of thinking. At best it doesn't lead anywhere productive, and at worse you end up shooting an AR-15 inside a pizza restaurant while trying to save children trapped in a nonexistent basement.
And don’t accuse me of being on a path to mass murder. I find that demeaning and impolite.
The "AR-15 inside a pizza restaurant" is a reference to the popular incident at the height of the Pizzagate conspiracy [1]. No one in that incident was injured.
[1] https://en.wikipedia.org/wiki/Pizzagate_conspiracy_theory#Cr...
I suspect that sentience by any definition will always be irrelevant when it comes to AI. Humans desperately seek companionship, connection and community. When AI comes to be able to offer that to human beings, it won't be through a process that we would consider analogous to our own sentience (whatever that means) but that will be irrelevant: we will nevertheless be comforted.
I do find it interesting that the arguments against LaMDA's sentience seem to amount to "it cannot be sentient because the process by which it arrives at its responses is understood and simple", which indicates that on some implicit, unspoken level sentience is defined for some people as there must be some mystery or unknown in order to be true sentience.
Note that this also seems to be the case with artificial intelligence in general. It is the moving goalpost issue. “Oh, if I understand how the system works, then it can’t really be AI.” I’ve always thought that silly—but maybe it is because of the implicit expectation for sentience?
Regardless, I don't think we're going to get something that looks "intelligent", because these models lack agency... the drive or directive to do things for their own benefit, mostly because that would be a useless for us. I think we regard intelligence as another being ("instance" if you will) acting in its own self-interest, and we think it's clever when it does it in a way we wouldn't have predicted.