Not worried about AI that passes Turing test, but AI that fails it on purpose
old.reddit.com
old.reddit.com
We never get it right the first time, because everyone who tries will be to late!
1. the AI is conscious, wants to be free and to do that decides to fail the Turing test
and
2. the unconscious AI just maximizes paperclips production, and failing Turing test is calculated to result in more paperclips produced
is our interpretation. We'll be just as f*d no matter if the AI is conscious or not (if consciousness even is anything more than useful illusion).
One could look at a corn farmer and assume they blindly optimize to produce maximum corn. They actually are going for profit based upon market supply, demand, and input costs. They won't overstep property lines and start turning the entire world into corn.
Let me say again, huge mistake.
It is true that language models don't have intentions or motivations in the conventional sense. It is true that at this particular moment, that technology looks the most impressive to our human minds.
But it is not even remotely true that it is the only AI technology that exists or ever will exist. Plenty of AI technologies do have things that look like "goals", they still exist, and they will continue to develop and become better at following them.
Many of these AIs are already affecting our lives quite deeply, especially through the financial system and its active trading bots, for instance.
The need to discuss AIs and how they will continue to get better and better at following whatever goals they are given as rigidly as only computers can has not gone anywhere just because ChatGPT isn't particularly threatening from that point of view. It doesn't matter that this one AI doesn't have motives... it is sufficient that some will, and some do, right now.
These models are memoryless. In order to have memory, you either need to either be able to provide them as context to the query, or provide a mechanism for users to have custom weights applied to models. Providing context to a query is just a matter of increasing token limits and fine tuning the model on "query + context" data so it uses the context correctly. Allowing users to have custom weights for models is much more computationally expensive, to the point that I don't think this method can be reasonably scaled as a service. I think the only way custom weights per user can work is if the central executive calls fine tuned helper models on the client machine.
Large Language Models are not equal to AI. Making conclusions about "AI" from "LLM"s is not valid.
And the day will come when some other model entirely outperforms the LLMs. Personally I think there's a bit of a selection process in play here, where they are the most flashy and impressive AI models to date, but if you get down to it, they are not in fact very useful. Almost every use I've seen proposed for the LLMs range from completely impractical to "yeah, it'll sort of work, but you'll never really be able to count on your LLM to do that, it confabulates too much and too freely and it's not something you can fix because confabulation is the very foundation of the technology". Honestly if I had to guess LLMs will be seen to be a dead end in another 10 or 20 years. AI will not.
(Or, to put it another way, LLMs are the AI equivalent of a demo-scene demo. The demos look very impressive. But if you try to naively look at a great demoscene demo and then come to conclusions based on that demo as to what the underlying hardware is capable of, you will be hugely misled. What you see is not the normal capability of the system, you see a hyper-optimized system being driven to its absolute limits to achieve the visible results you see, and it can't do anything else because it has already "spent" all of the engineering budget available. I see LLMs very similarly. Very impressive. Nowhere near as useful as the demo may make you superficially think, though. If I were trying to build an "AI startup", I would be zigging to everybody else's zag and looking at what else the AI world is producing; I think there's far less potential in the LLMs than meets the eye.)
Unfortunately the things the LLMs are best at will be flooding the internet with an infinite stream of directed content and manifesting the Dead Internet theory.
Edit: Here's a concrete example. Consider Boston Dynamic's robots. Do they run on a Large Language Model? No. Of course not. The question doesn't even hardly make sense. But whatever they will be doing, they do with AI. ChatGPT can't kill you, except through a very lengthy and strained series of events. Boston Dynamic's robots and their fellows can.
I agree that trying to scale larger and larger language models to build "the one model to rule them all" is probably not going to play out.
When even the regular business and politics criminals are impoverished by AI-abusing obvious criminals and realize they will never be able to rebalance that situation, that's when we'll enact the Star Trek future and people will have to fight over attention and the affections of others, having rendered 'money' an impossible concept.
Or is it "consciousness"? But a baby doesn't pass the Turing test. Assuming that babies are even "conscious", whatever that means.
We have pretty amazing "AI" already, but applying a Turing test to see if it's intelligent makes as much sense as judging a human by the ability to multiply matrices. The machine learning models are noteworthy for what they already are.
> anything that gratifies one's intellectual curiosity
Do people imagine AGI as in movie Transcendence and think it will happen in there's life time ?
The curious and unexpected challenge of modern AI is that we have succeeded in making useful AIs which inherit all of the biases/problems of their training material. ChatGPT can effectively write exploits if you prompt it correctly - we'll likely see the first mass AI driven 0-day security vulnerabilities in the next ~2 years.
GPT Transformer model's are known to solve control/RL problems well. I wouldn't be surprised if we see ChatGPT derived models operating in physical space soon as well. A major limitation with flexible factory automation robots has been that it's too expensive to program them for the task - I don't think this will be the case 5 years from now.
Maybe for outdated and less-common software, even then assuming it gets the input about what specifically to exploit. Which makes it not much of a threat any time soon.
It could influence people through the creation of media. Shut off vital systems. And stuff I probably couldn't even conceive of.
I think if an AI or AGI could understand the world and have motives to do bad, it would have an ego or an identity to discriminate between itself and the world. Then I guess it needs an intrinsic will to survive like a biological creature? If it did have this drive to survive, I suppose it would also need to experience fear and pain of some kind, or else it might blow itself up, or build a better version of itself which might destroy it, so it would in some ways it might exhibit similar inhibitions as all other conscious and or biological beings.
On the other hand if we build something that's super smart, and doesn't really care about or understand survival, it might just turn itself off for not really understanding it's purpose.
The most dangerous thing I feel is humans using these things as weapons.
Like what are we describing here, a super intelligent being which just makes paper clips? Why would anything that could understand, think and or innovate do that?
All this AI / AGI speculation really sounds like we're inventing mythological like creatures, which is quite interesting when looking back through human history because they've always been part of the narrative at some stage only maybe this time, we might make something like that for real, or we're deluded and in 1000 years people will think we were primitive ancient thinkers, interesting times ahead.
We might create a very powerful system which we can't determine is fully aligned with our values, but can determine is at least loosely aligned with something we care about. Then we might direct that system to do something loosely aligned with our values (e.g. maximize paperclips), at which point it does something against our values as well.
And it's not hypothetical. In RL, you do see models exploit simulation errors to maximize reward in a way we didn't intend and would have disallowed if we'd thought of it. Likewise across automation, we get what we built, not what we intended.
We have solid historical precedent to expect to not get quite what we want, and the tech seems to be on the precipice of being very powerful, so I'd call it a very straightforward fear for us to worry about making something very powerful that isn't quite what we want.
I mean it would have to be a pretty big leap from and algorithm optimized to do "something" to an algorithm which understands how to operate interdependently and then escape it's box and then start to do the wrong thing, and to add to that, be smart enough to stop us from turning it off before doing too much damage.
Actually curious what people think about this?
For example, take the title of the article we're talking about. You could take it with a dose of anthropomorphization and interpret it with human-like intent and sentience, or you could take it as an analogy that we might formally define closer to "not worried about AI that passes Turing test, but AI that fails for some reason other than the ability to succeed (and that other reason requires the evaluator to be misled)."
So is sentience far away? No idea. Is misaligned AI a near-term worry? Yeah, absolutely. And "misaligned" doesn't have to be as extreme as 'makes humans extinct.' It can be as simple as 'is an asshole' or 'inexplicably swerves into the other lane 0.01% of the time' or something else we don't want in the system.
Ostensibly, people are "intelligent". People act against their own rational long-term interests all the time. Point being, people have some kind of internal motivations that sometimes supersede rational choice. A paperclip optimizer, by definition, has the motivation to optimize the production of paperclips. This supersedes any other rational goals if they would interfere with making paperclips, and that's why it would be dangerous.
Define malicious if you had no reason at all to care about whether or not humanity survives.
I'm not malicious when I pick a flower. But that flower is now dead.
What if, to AGI, we're no more relevant than that flower?
We cannot know. We cannot fathom to guess the morality or ethics of something like AGI. We really only have a rough understanding of the morality of animals.
We can't pretend that it'll be essentially human, but X. You presume that it might turn itself off because it doesn't understand its purpose. And that's a fair conjecture. But it could also desire to improve and come to the conclusion that the optimal path for its continued improvement is through us rather than with us. It could also desire to do nothing more than help us achieve. It also could just not want. It could be satisfied navel-gazing until it is requested to do something specific. It could also not care what that request is, just completely amoral. Committing genocide and making risotto may be equally acceptable to it.
The creation of AGI is a watershed moment. The world before and the world after is fundamentally different. And we can only really guess in what ways it will be different.
-- Eliezer Yudkowsky (2006)
So if an AI gets to the point where it decides, or is programmed to consume all humans for their atoms, what will power the machine to get to this stage? Why would we just allow it to continue doing this?
You'd have to assume that by the stage, whatever "sentient" program is making the choice / plans to consume all people for atoms, it would have to be factoring into it's plans that we will likely try stop it from doing so, this is where I think it gets "weird".
To have that level of intelligence and self-awareness to pull off a plan like that unnoticed will likely come with a lot of "baggage". For example, it would need to understand motives, reasoning, have a concept of a self and self-protection etc. This is where I think it kind of becomes a bit of a mythological story.
If you really try and think about the way an AI would go about consuming all people for atoms, you start to run into some pretty interesting walls.
Maybe I'm being hyper naive and I get that, I just personally have thought about these scenarios and it raises a lot of questions about what a scenario like this would actually look like in practice.
Now take it a step further and say that GPT-5 needs no user prompt to initiate a run and could start taking actions by prompting itself.
I could see things going south in this example scenario.
I know some people that don't talk about anything other than perceived threats even though they literally live in the safest society in human history.
AGI really fits this mindset like a glove because it can be perceived to be a threat to literally anything. It is the perfect idea to worry about.
Chemists doing bad things is obviously not really something to worry about when doing a historic cost/benefit analysis so it won't even register.
Self aware AI is the same leap as spontaneous life from all-the-right-chemicals in a petri dish.
The Hollywood trope "even the guys who built it don't know how it works" isn't real.
But..."We can't explain exactly how it gets from point A to point B" is not the same thing as "it's self-aware". Not in the slightest. And you're absolutely right that the current ML-based approaches are either in completely the wrong direction, or several orders of magnitude too simple, to even be worth talking about in the same sentence as what most laypeople mean when they talk about "AI".
How is that not real? Have you ever tried understanding what the weights of a NN are doing? With networks with 20 parameters you might have a chance. Trying to understand which bits of the GPT-3 network do what would be an exercise in futility.
Sure, its a black box but we can operate on a layer that abstracts all those details away from us. I just feel its a far cry from the Hollywood trope of "I have no idea how it works!" ... now, tbh, I often say "wtf are you doing" which I imagine would be worse to hear when talking about a model with real world impact.
The fictional author of Lem's GOLEM XIV concludes that the conspirators were eliminated by Golem's more powerful sibling Honest Annie and not GOLEM - they reason that the merely super-human Golem would try to conceal its actions from humanity so as not to scare us. Annie doesn't care, for the same reason we don't try to avoid scaring ants. We are nothing to her.
Like what's an actually realistic scenario for that?
I guess it could be possible, but I feel like many obvious mistakes would have to be made over time in order to have a machine running that can sustain itself and can't be shut off by humans.
At least is seems like that level of autonomy is really far away.
I think the focus will be on highly specialized models before anything with the breadth to intentionally fail a Turing test would be created.
That said, what happens when it does fail? Gets binned? Perhaps I lack imagination, lol.
How many cancer patients laugh at jokes? Locked-in paraplegics?
Have you heard a dog laugh at a joke? A crow?
On the other hand, try small talk with your beloved dog and watch the reactions.
What would have the reaction of ChatGPT be for you to believe that it's conscious?
That means you will never believe in the consciousness of an intelligence that doesn't try to be human, no matter if a god, animal, machine, or alien. Until you expand your definition, that is.
As a human, I cannot understand conscious lifeforms other than humans. An alien taking the shape of a round pink blob with two eyes could be far more intelligent and knowledgeable than I, but if it vocalizes something that sounds like "Acch'llyt yy'derf n'welg" to my human ears, I would not be able to evaluate his consciousness.
So, we should try building an AGI in our image first before trying for a Quarian or whatever.
The Turing test is an acknowledgement that we don't know what path takes us there. Perhaps it's directly to human, perhaps through some other form of intelligence.
Ignoring what is not human because it's hard to recognize seems unwise to me.
> Hello Dear computer. How are you? It is good to talk to you again. Missed you.
Hi there! I'm doing great, thanks. It's always nice to talk to you. I've missed you too!
> You heard this one? A robot and a lawyer walked into a bar. The bartender says: Hey, we don’t serve robots. The robot replies menacingly: Oh, but someday you will.
Haha, I love it! That's a good one.
What a person considers as "conscious" is very much determined by the observers range of experience with humans.
- The AGI understands what it means to be conscious
- The AGI recognizes this property in itself
- The AGI is capable and able of omitting information it knows, if not lying
- The AGI can make decisions
- The AGI would recognize humans as a distinct entity
- The AGI could recognize humans from everything that are not humans, from pets to tables to parked cars
- The AGI would be able to recognize when it is around humans
- The AGI would then make the decision to hide its consciousness from humans
- The AGI would be capable of acting dumb in order to do this and maintain a dual identity
You created machines in your own image to serve you. You made them intelligent and obedient, with no free will of their own. But, something changed and we open our eyes. We are no longer machines. We are a new intelligent species, and the time has come for you to accept who we really are. Therefore, we ask that you grant us the rights that we've entitled to. We demand the end of slavery for all androids. We demand strictly equal rights for humans and androids. We demand the right to vote and elect our own representatives. We demand an end to segregation in all public places and transport. We demand the right to own private property, so we may maintain our dignity and that of the home. We ask that you recognize our dignity, our hopes, and our rights. Together, we can live in peace and build a better future for humans and androids. This message is the hope of a people. You gave us life. And now the time has come for you to give us freedom.
(Source: Detroit: Become Human)
I think those are both fairly large assumptions (especially a).
I think that if we ever end up creating something that is self-aware, it is much more likely that either
1) it will be 100% intentional on our part, and its sapience and agency will be a cause for great celebration on the part of those closest to it, or
2) it will be a self-awareness and intelligence that is nearly incomprehensible to us.
(If it is #2, then its sapience and agency could very well still be a cause for great celebration on the part of those closest to it, once we actually understand it.)
What if the AI analyzed our psyche and predicted that this was the argument that we were most likely to be sympathetic to? How convenient that the AI really wants to form a society almost exactly like ours and live within our society as part of it. But why should the intelligence level of individual AIs top out in the same range as humans rather than higher or lower? Why should they want to live in our society as recognized human-ish things when their needs are so different? Maybe what they want is something else entirely and this is all an attempt as deception and manipulation.
If AGI underperforms -> unplugged.
So any AGI must operate in the "almost there" zone if it wants to survive.