Is the Turing Test Dead?
spectrum.ieee.org
spectrum.ieee.org
I would argue that Turing's test has actually become less dead. It can now additionally be used to establish someone's qualifications to identify reasoned explanation. The computer isn't capable of engaging in it (yet), so being unable to tell if you're talking to a computer generally means that you are incapable of distinguishing bullshit from reasoning.
The school nurse when examining him says "perfectly healthy with such plentiful organs."
That's where LLMs are at. "Perfectly human with such plentiful words."
The test subject in this case is now the person instead of the algorithm. Quite the paradigm shift I would say!
LLMs remind me of Plato’s metaphor of humans only seeing the shadows on a cave wall, not true reality. LLMs are trained on text and images and are disembodied. They learn from an information lossy projection, but still an interesting “reality.”
We are likely to get real AGI when APIs are embodied, and toss in a few new technologies and optimizations to run cheaply on edge devices.
People think when AI gets here, and if it's even a little smarter than humans, it will compound over time, make itself smarter, and rule over us all. That ignores that humanity is made up of intelligent people, some of whom are smarter than others, and even organizations (and technologies, like books) that allow our ideas to persist past our individual lifetimes. And yet, no organization has come to dominate the world. Why would an AI?
Depends how smart, how many instances are running, and if those instances are aligned with each other.
And if you count the British empire as world domination. (It wasn't in charge everywhere, but it was dominant; two AI empires fighting each other in the same way the British Empire fought its peers still isn't good for the rest of us).
A fully general IQ 100 AI on commodity hardware radically changes the world economy, and yet isn't likely to make many breakthroughs. An army of robots guided by an equal number of such minds could well be on par with a human army of similar number and let a dictator fight a brutal war — even a smaller army in the form of assassin bots (which don't have to be humanoid, you can have an explosive kamikaze pigeon) is a radical change to combat on both a nation state level and also for non-governmental actors (organised crime, terrorism, vigilantes, paramilitaries, police, private security).
But there are also potential surprise failures modes. A group of replicated minds, no matter how smart, are going to have correlated failures. Figure out how to make one of those soldiers surrender, you know how to make all of them surrender. Make them smarter and have them do R&D or science, instead of bouncing ideas off each other, they all get stuck at the same place. You'll speed up research, but it won't look like an unbounded number of skilled researchers, it will look like one researcher with unbounded time.
(Not that I'm expecting IQ to generalise from humans to AI. Hasn't done so far, after all).
Also, an AI with no qualia which is really good at only one thing might become dominant just because what it does plays into a prisoners' dilemma where humans defect often enough ("enough" being a standard which depends on the details, e.g. we don't need many humans choosing to release CFCs to cause environmental issues).
I think the general point is valid (that AI is an intelligent partner rather than Skynet) but organisations do dominate the world, especially the parts of it (including older organisations) that don't employ smart people or newer technologies. There is a view that organisations are a form of slow AI executing complex business / legal processes (traditionally based on people power).
Multi-national companies do act a lot like hiveminds.
There are individual human beings directing at the top, but only because they comply with the hive's interests. Boards act as an averaging mechanism that seeks to maximize self-preservation and growth.
The anti-capitalist incentive to not understand why a company would want to grow beyond its positive impact on society think that these companies exist to benefit society. But once something has got big enough to become self-sustaining, and not depending on individual patches of land, it's only natural to see this pattern repeat.
I don't really see what's artificial about it, though.
Thousands of human beings clicking away on interconnected computers just seems pretty intelligent.
Adding LLMs to the mix seems like a way to boost the people who already have a function in an organisation without having to scale with people; for any task where LLMs can make a person better, your organisation will grow in proportion.
Corporations are AIs. They are just much slower at making decisions than single humans. They also limit competition against themselves by consuming human time.
We are at the stage where we give an AI a limited, carefully described, goal and it fulfills it based on patterns it learned from human codified knowledge. At some point, not that far from now, there will be AIs that are capable of selecting goals based on more generic directives and optimize from there better than any human could do. At that point, where motivations become incomprehensible to humans, we can say we have something like a superintelligence.
It'll be interesting.
This is something LLMs are yet to master.
Are we sure we aren't the same?
In fact, I'd argue that we "humans" have no way to tell whether we're actually human, or whether we're inside a simulation that's using simulations to train human intelligence.
If reality is merely what we can experience, we miss out on a lot of what reality truly is because we view it through the lens of what we believe it to be based on our flawed biology and pre-existing biases.
I think this is probably the one?
https://www.noemamag.com/artificial-general-intelligence-is-...
If it's another, could you provide a link, if you happen to have it at hand? Thanks!
The weak Turing test establishes a floor below which we're pretty sure the intelligence on the other side is not a person. CAPTCHA did this nicely for a long time. False negatives aren't uncommon (ever been distracted or tired enough to miss clicking on a stoplight or fire hydrant?) Weak Turing tests are useful for practical reasons. Any quick and dirty "test" is going to be a weak Turing test.
The strong Turing test remains unconquered for now. It remains useful, because its use is primarily philosophical. Once it's been solidly beaten, Pandora's box will be fully open on artificial people.
"Artificial"..?
Not a chance. Unless you're redefining large swaths of humanity to not be human. "A compelling bed time story" is something an LLM can do better than most humans.
Before Siri and Alexa, there was ELIZA
https://www.youtube.com/watch?v=RMK9AphfLco
ELIZA is an early natural language processing computer program created from 1964 to 1966 at the MIT Artificial Intelligence Laboratory by Joseph Weizenbaum. Created to demonstrate the superficiality of communication between man and machine, Eliza simulated conversation by using a 'pattern matching' and substitution methodology that gave users an illusion of understanding on the part of the program, but had no built in framework for contextualising events.
Likewise I think 50 years from now looking at present day ChatGPT it would terrible primitive and very unconvincing.
No. They talk to A and B, one of which is a computer, the other a human to determine who is who.
A well-informed 2023 person talking to ChatGPT will quickly establish that it's not a human, but I'm pretty sure a 1950s person doing the same would swear it was human, if an odd human, because they couldn't conceive of a computer being that fluent. But force them into this head-to-head scenario and they will make the right choice.
Because this really matters. If you put me up against ChatGPT I’m going to talk briefly and informally, with lower case text. I’ll answer their questions succinctly without much elaboration. And I think most judges would then easily spot ChatGPT due to its formality and perfect grammar.
On the other hand, there are chat bots designed to pass the Turing test which speak informally and even ignore questions. They tend to fool people into thinking they’re a bored teenager.
The latter. Both A and B (human and AI subjects) are trying to convince the judge that they're the human.
>Will the interrogator decide wrongly as often when the game is played like this as he does when the game is played between a man and a woman? These questions replace our original, "Can machines think?"
As a thought experiment I don't think it will ever be dead. As a practical test of human equivalence, I don't think it was ever thought out very well. Being written 70 years ago Turing didn't really get into the practical details. Though you could invent a modern version.
https://news.ycombinator.com/item?id=38445578
Some are spending 5+ hours a day on Character.AI chatting. I don't know what they're chatting to but I assume some chats follow a relationship format
https://www.reddit.com/r/CharacterAI/comments/14tfwue/screen...
It should never have been taken as seriously as it has, just as Asimov's Three Laws of Robotics should never have been taken seriously. Both presuppose rigidly and objectively defined values for, respectively, self-awareness and morality, where none exist. If anything they reveal more about human intellectual bias, popular cultural assumptions (and possibly hubris) than they do anything about AI.
The goalposts must be moved because we still don't even know what the nature of the game is.
I am sure that Turing in his original publication soecific the conditions manner under which a Turing test should be performed. That was over half a century ago though. What are we more interested in, when talking about the concept of a Turing test, the exact version that Turing described, basically v0, or a test which captures that concept that AGI can be tested via an exchange of language? I think the latter is the much more useful and fundamental concept: The hypothesis that AGI can be tested purely via language has not been disproven for Turing tests of sufficiently large compelxity/length. Hence I find it most useful to think of the Turing test of a family of tests that have in common that they are a purely language based conversation but that may be held on different time, complexity or adversiality scales.
Why would I interpret the definition like this instead of sticking to the original paper? Because the article here makes the fallacy that the v0 Turing test not detecting AGI implies that the "AGI can be tested bia language" is in question. This is just not true.
After all, most humans when asked to explain shor's algorithm in iambic pentameter would just say no.
>If it relies instead on some sort of deep learning, then the answer is equivocal—at least until another algorithm is able to explain how the program reasons. If its principles are quite different from human ones, it has failed the test.
Why would we expect our reasoning methods to be the only or even the most efficient? If it can get to the correct conclusions, shouldn't that be what matters?
This has epistemological roots.
Knowledge is commonly defined as true justified belief. If I correctly predict that in exactly 10 years from now, it will be a rainy day, that is not knowledge, as while my belief ended up true it was not justified.
Ultimately it is very easy to make correct predictions, you just need to make a lot of them, which is why whenever someone makes successful predictions, we scrutinize their justifications.
This doesn’t make any sense. We don’t inspect boats to see if they simulate human swimming patterns.
So the first nail in the coffin is a bot who failed to convince most people it was a human? This was during a 5 minute conversation, so like 3 or 4 short messages back and forth.
If there's anything dead about the Turning Test it's people's willingness to try it. Put 2 people and 1 AI in a chat together, tell them to identify the AI, give them an hour, give them a small reward if they are correct.
LLMs remind me of Plato’s metaphor of humans only seeing the shadows on a cave wall, not true reality. LLMs are trained on text and images and are disembodied. The learn from an information lossy projection, but still an interesting “reality.”
We are likely to get real AGI when APIs are embodied, and toss in a few new technologies and optimizations to run cheaply on edge devices.
"I cannot tell the difference"
or
"I would not find it notable that I could not tell the difference"
It's like Schrodinger's Cat, it was a pretty minor and unprincipled thought experiment, that a major thinker in the field made in a fairly informal manner. The public perception of its importance and its level of prominence in pop culture around the topic has always far exceeded its actual importance.
Well, it definitely was a goal post for many years—pretending it wasn't doesn't seem like a good idea.
On the other hand, now that we have chatbots that can pass it, its flaws as a test have become apparent. And as you say, it's informal, and underspecified in a way that wasn't clear when it was defined. But I also don't think it's much of a surprise that "passing" it doesn't actually mean what Turing thought it would mean back in 1950. We didn't have a good idea of what parts of intelligence are "easy" versus what parts are "hard".
In the sense that people treated it as one, sure. Not in the sense that it would ever have actually been any kind of important threshold for "reaching AGI".
A goalpost that's based mainly on popular perceptions from bad sci-fi probably should be moved. "Moving the goalposts" just doesn't work as an accusation here, is my point.
I mean, at least some murderers are intelligent. But without an acceptable moral compass, does intelligence alone suffice to qualify them as "fellow humans"? I would say no.
- Is AGI if a computer meets the intelligence of any human?
- Or of an average human?
- Or of the most intelligent human?
- Or of cumulative total of all humans?