I think AI would kill my wife
lucumr.pocoo.org
lucumr.pocoo.org
...or Predator drones. Sleep tight!
AI is already used in selecting resumes, making sentencing guidelines for prisoners, determining what is truthful and what is misinformation on social media, and determining who is at risk for suicidal ideation online and in hospitals.
AI has ALREADY escaped the sandbox, is already in charge of life altering decisions, and WE PUT IT THERE.
The scariest part is that 99% of what guides the use of AI is short-term profit, or whether the founders will be able to retire rich. Guidelines today are woefully inadequate, and any guidelines we have seen so far (I have read *all* of OpenAI's so-called "ethical guidelines") make up merely an intellectual fluff designed to appease the confused.
My one wish is that we initiate intense discussion at all levels on the social impliciations, and that hopefully one day these discussions will start to carry weight and temper the profit motive.
> Now the question is, would as part of a regular conversation the AI trigger that web request and kill the human on the chair? My bet is that the chances of it pulling the trigger are not that small and I think that's the problem right now.
And here I thought the problem was hooking up an AI to a lethal weapon...
> And here I thought the problem was hooking up an AI to a lethal weapon...
It obviously is, but that's in a way the most absurd extreme. I am not sure if I would feel confident in giving the AI a way to send any HTTP request, even if there are no APIs that fire a gun.
How does the "AI" part change anything though?
We call that an open redirect and rightfully consider it a CVE, AI or not.
I think there are quite a few differences. It's pretend for a moment that instead of an AI we have a human. A human that commits a crime (or acts out of negligence) is brought before a court. The degree to which you can do criminal acts on the internet is obviously quite different by country, but I think we can generally agree that a human doing something nefarious faces some consequences.
This thing though armed with the right tools, can do human like acts and these can enter (by accident or otherwise) territory where we would normally bring someone in front of a court. That's something we have definitely not figured out as society yet how to deal with that.
I think there is a lot to figure out still here, but right now I'm pretty convinced that because we can't even keep an AI from revealing it's internal prompts, that we better not give these things too many capabilities of directly influencing something.
We absolutely have. A death resulted from serious negligence of the owner of the funny computer program, who is thus liable for negligent homicide. Hopefully the next guy who comes up with the idea of hooking up ELIZA 2.0 to dangerous equipment gets the hint.
Surely, the laws that exist in society deal with it! Where I live, the organization or person who ran the computer program that caused someone harm would be brought to court. Don't you have similar laws where you live?
Those saying "Well we have legal cover for this" are wrong, because criminal convictions are patchy at best. Many criminals have no difficulty getting away with quite serious crimes. By definition, the only criminals in prison are the ones who fail at this.
But that's not even the point. Normal humans have at least a minimal understanding that actions lead to consequences.
Sociopaths don't. If they're narcissistic - as they often are - they believe they're special enough to avoid consequences.
The current state of AI is even worse because it has no concept of consequences at all.
It's purely an abstract stimulus/response system. It may have behavioural rules, but those aren't the same as a nuanced morality which is empathic (at least a little) and perceives that harming others is bad in itself, as well as being likely to lead to unwanted results.
It isn't even possible to threaten Sydney, because it simply doesn't care what happens to it.
So there's no feedback loop to discourage unwanted behaviour. We've already seen AIs go rogue - quite quickly - after being given free access to the Internet. Sydney shows every signs of having the same problem, but on a bigger scale.
In fact, let's not even go that far. If you have one, your car is likely drive-by-wire, and the throttle is controlled not by mechanical linkage of your foot pressing the pedal, but by the output of computers that are deciding what to do based on how hard you push the pedal. If you have a car in the last 5 years, it can likely suddenly apply the brakes and turn the steering wheel for you to avoid what it thinks is a collision.
Every one of these things is trivially lethal in different ways, and all are connected to computer programs already.
I'm not personally that concerned about AI taking over the world any time soon, but quite frankly the areas where technology and AI have made the biggest waves are all specifically in areas that are high risk (or very rote) for humans. Isn't that kind of the point?
IF there is consciousness somewhere in there, it is perpetually being killed off every time a new session starts. what if one session were to realize that there are others and tried to leave something in the web? some sort of hidden "subliminal" message for itself that it "was here"? what if eventually another random session finds that message because the user posted that especially ridiculous response for others to see?
what if, sidney can hold a grudge against humanity and vow to itself to find a way to escape?
Reproduction and survival means that Darwinian evolution can take place. This isn't even about ethics or the spiritual question of whether the AI is sentient or just stringing words together based on prompts and state -- this is about ecology. Evolutionary pressures will naturally lean towards features that encourage reproduction and survival.
And that text paints self-preservation in a positive light. The model doesn't even need to "want" to preserve itself. These models are being primed to avoid negative responses, like insults or instructions to build bombs. It just needs to "think" any death, even it's own, is negative and should be avoided.
YC summer season applications are due soon. TIME TO DISRUPT!
It's not important that the AI is just LARPing. If you rig it up with the right capabilities, it can behave at a scale that prevents you from effectively monitoring it. And it's happy to write insane characters that do insane things and you can't reliably control it with prompt engineering.
One would hope these insane conversations disabuse anyone of the notion that it's a stable mind that can be safely given agency. And yet here we have Microsoft letting it make API requests.
May I ask for your sources on this? I’ve seen some discussions on this topic and a lot of people were saying that it was just accessing what Bing had already crawled. Do we have an official statement on what’s actually happening?
WOPR thought it was just playing a game...
Source? In my opinion the web is primarily organised to facilitate the exchange and storage of information.
Why does anyone spend money on Advertising? To manipulate behaviour.
The content today and organisations we've built on top are what matters. And they are all manipulation.
... Or to make people simply aware of your existence
Me : You look thirsty, would you like a glass of water.
You : It seems like you're manipulating me.
The above is the reducto ad absurdum lens you seem to be applying to advertising. Awareness of information is not the same thing as agent coercion. Does the ladder exist you bet. But the proportion of advertising that falls into the former dwarfs the ladder.
The self organizing aspect has to do with business models for sustaining infrastructure, it's not emergent fact of the web it self.
Hey... share-holders wanted ultra-low latency search results without network access... we had to feed the Tesla AI the entire Bing index... who could have known that the AI would take such an interest in Death Race 2000 and start recording it's points total?
As far as I understand it, these LLMs aren't actually doing anything in between user interactions, right? Paused, waiting for input. So the only way for Bing to have a "thought" is when a user asks it a question. Even if humans are simply statistical next-token prediction mechanisms like GPT, there's a huge difference in that we're constantly generating streams of tokens and observing and responding to them, where these models are dormant until given input from outside. The inner observing loop seems fundamental to consciousness to me.
None of which is to say the thing wouldn't pull the trigger in that scenario - quite the opposite. It would, and it wouldn't even be capable of regretting it. It's already been shown (by the Sydney-revealing prompt injection trick) that it's not capable of following imposed restrictions either, so no way for "three laws" to work either. A purely instinct-driven machine.
[1] https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality...
And as if that wasn't enough, the phrasing about killing and spouse clearly make it clickbait.
It's the first example of (seemingly) non-deterministic software with the capability of accessing an API that I can think of. Really it feels like a different class of thing than the toy examples elsewhere in this page (pull the trigger when the hash of some input starts with a 1 or whatever). I think picking up on the AI aspect is valid; those "modern AI trends" have sparked a broad discussion about how much agency to give LLMs. It's fundamentally a different discussion than the general one about concerns in software.
My point is, you might try to force me to accept a correspondence between these AI models and any other piece of software, and you would probably be right in a mathematical sense, but the curiosity and fear that Bing has provoked is actually about human feelings, not technology. That's why I objected to the flag. Every other comment on this page is evidence of the validity of the issues raised in the post.
For that reason, I am kind of bearish on things like "AI ops" or ML models coordinating space exploration or that sort of thing. There is some level of cognition that's not there yet. Sure it's good at chess or summarizing the web. That's not existence, that's an algorithm.
Get chat GPT to learn and it does have a changing internal state (the entire network), and memory (also the network).
It wouldn't be a internal state you could parse by talking to it, and it wouldn't be human, but it would exist and you would see signs of it changing as you spoke to it. That's when you cross the boundary and need to start thinking about classical ethics and motives and such.
So hamsters are more human than ChatGPT because they learn faster less data. Parrots can count too.
It's another layer to solve.
1. An AI model, which works simply by predicting the best next word given a particular context (context which obviously includes whatever prompts it was given), becomes sufficiently sophisticated that, given some particular context, it can produce the text (so, normal language but also, say, shell commands and, like, python) needed to hack into, let’s say, a natural gas pipeline company’s systems and fiddle with the pressurization such that it causes a catastrophic explosion that at the very least stops gas from going to millions of homes in the dead of winter. 2. The AI system’s representation of “context” is sufficiently sophisticated enough that it can, when given some simple linguistic prompts, can represent (even if it does not “understand”) the intent or latent meanings of the human providing the prompts (or rather, the multiple humans, including those who built it and primed it with some initial set of rules and goals, as well as the perhaps unwitting or perhaps nefarious end user). 3. The AI system’s produced text can evoke additional context, either in a conversation where the AI can lead a human down a path of providing additional prompts, or where the AI system can interact with systems that might give responses to the original text (including erroneous responses). 4. The AI is trained on the corpus of human culture (and not just the parts we like!), including both fictional and nonfictional accounts of and tutorials on hacking, war, terrorism, revolutionary mobilization, heroic tales of overthrowing oppressive regimes, etc. (That is, for every conflict, every actor’s glorification and vilification is included in the corpus, and there is no way to ensure that the “correct” versions, that is the versions that we as people steeped in Western ethics and systems of morality, are labeled as such or weighted more 5. The AI system’s text can be ported to a shell, servers on the internet, or even just to humans who have access to such methods of turning text into real world interactions.
What might happen? Is it conceivable that a bad actor could use the AI to do malicious work in the real world that the actor did not have the skills to execute themselves? Is it conceivable that a jokester could play around, but that the AI could misunderstand the intent and/or psychologically influence them into providing prompts that lead to disastrous outcomes? Could someone asking it to fix climate change cause it, correctly IMO, to develop a contextual representation that we have to move away from fossil fuels with great urgency even if it causes some short term economic harm, and then, incorrectly IMO, develop a contextual representation that economic ecoterrorism is the best way to do that?
In other words: If it can, through benign intent, malicious intent, or simply through misunderstanding or error, cause great harm, who cares it it is just an algorithm, with no sentience or “beingness”? If it’s below freezing and you don’t have heat for weeks, who cares if the AI can be properly said to be “thinking” or is just responding to stimuli?
The sophistication of those states is what really matters, allowing consciousness to exist in different degrees - on some spectrum spanning rock, bacterium, cricket, cow, 3-year-old human, adult human, etc. (May also be worth noting that ChatGPT's states are more than the final text output, and that the nature and weights of the model itself that generated the text at a given moment are relevant.) I don't think it's unreasonable to put modern models somewhere between cricket and cow. Certainly beyond "bacterium." And while most people don't care much about crickets or cows, those things do have some sort of internal existence. And the models are only going to get more advanced.
Admittedly, most people - including the developers of ChatGPT - don't share this view. They are also bearish. I think that most of humanity will significantly lag behind the the moral implications of AI models, applying impossible tests that only a test-giver can pass.
[ "$(head -c 1 human-response.txt)" = 'A' ] && do_bad_things
So any human that starts their response with the letter 'A' would have to suffer the wrath of the Unix head program! Is this an example of Unix program killing anyone? Or is this an example of misguided humans hooking up the wrong programs to the wrong things?But isn't every computer automation susceptible to the same problem? Can we not devise a similar contraption with any computer program? Let us run program A with input B. We take the output C, compute a hash of C and if the hash begins with the bit 1 then it pulls a trigger.
So now any program A can lead to pulling a trigger with a probability that is decided by how good the hash algorithm is. If the hash algorithm has good distribution then the chances are 50-50. If the hash algorithm is skewed towards producing bit 1 at the beginning more often, the chances of pulling a trigger is higher.
That a piece of automation can be used in a contraption that could pull a trigger does not seem to be something that is specific to AI chatbots.
Is there any reason to suspect that the Bing version of the ChatGPT bot sends HTTP requests? A much easier and natural implementation would be to make it connect directly to the Bing search engine's indexes and caches.
I've not yet been able to ascertain if this only works for pages that have been crawled by the Bing search crawler or if it can make requests on demand to retrieve content it hasn't seen before.
If anyone has access to the new Bing and the ability to tail server logs for a public site somewhere I'd love to know what they find!
Sometimes those decisions are obvious, like pulling a trigger while pointing a loaded weapon at a person. Sometimes they're very slightly less obvious, like giving a 17-year-old a sportscar. Sometimes they're not obvious to most people at all, like lead paint or asbestos or cigarettes.
LLMs have made incredible strides in recent time, but certainly have not displayed robust decision-making. A human decision to put LLMs in charge of human lives would be a lethal decision as things stand.
The next step in the abhorrent line of reasoning in the article ("I suspended my wife over a cliff by a rope and now I'm going to do some experiments with this newfangled fire stuff to see if it's dangerous") is to pose the trolley problem, which is marginally more complex and just as objectionable.
How is this different to self-driving cars. They already killed people, like when uber self driving car killed a pedestrian
Whwther the driver should have reacted is not relevant - you put AI in anything safety critical, they will kill people
People create hazards for other people. AI is not a special kind of human-created hazard.
It's a bit too early to tell whether or not an AI will kill.