Turing Test Success
reading.ac.uk
reading.ac.uk
>It will simplify matters for the reader if I explain first my own beliefs in the matter. Consider first the more accurate form of the question. I believe that in about fifty years' time it will be possible, to programme computers, with a storage capacity of about 109, to make them play the imitation game so well that an average interrogator will not have more than 70 per cent chance of making the right identification after five minutes of questioning.
COMPUTING MACHINERY AND INTELLIGENCE
— A. M. Turing
10^9 bits is only (check my math) about 120 MBytes.
So Turing was perhaps a bit optimistic there?
A more correct test (which admittedly doesn't cover this issue) would be to give each judge a conversation with one human and one computer, and for them to say which one they believe is the human.
I always assumed this was exactly what the Turing test was about. Guess I was wrong.
The Turing test is not a test. Instead, it is an operational definition of "intelligence", a very slight formalization of the idea that something is intelligent if it seems to be intelligent.
As a test, it obviously has to have some kind of limits like this "competition", but as soon as you put limits on it then it stops being useful and becomes both gameable and meaningless. The Turing test has already been passed, long ago, if you have limits suitable to the Doctor or Parry.
Exactly. It's a straightforward formulation of what a strong AI would be capable of. It makes no sense to have a restricted Turing test that can be passed by a useless chatbot. It means absolutely nothing.
It's a way of having a better metric than pass/fail. This is the real problem with AI. Everyone expects that AI research should go: you put a bunch of programmers in a room, and they work for a few years and build an AI.
The best intelligence production we have is human child-rearing. This process has always taken 15-20 years, and is backed by millennia of research. Assume you have a digital computer capable of human-like thought. Without an example of a computer capable of learning faster, it stands to reason that raising the computer into an AI capable of conversation should take 15-20 years.
Of course, one thinks there would be a way to do it faster, but on the first go?
[1] Although usage may be changing. Pity.
I studied AI shortly after the "AI Winter", which is where I got my definition. Strong AI---working towards a general intelligence---was strongly out of favor, especially with funding agencies. It still remains so (but see Watson). But solving limited problems heuristically (or statistically) that are not otherwise algorithmically tractable (a loose translation of "would appear to require general intelligence") has always been a fertile field.
Turing's argument, which is a philosophical argument, is not meaningful if you put any limits on it---time, topic, behavior (really, Parry is better than the Doctor), which is why it is better thought of as a thought experiment. If you have limits such that it can be gamed, then yes, it is fair to say "you only need to simulate intelligence well enough to fool the judge." Which makes it uninteresting.
But the question is, if you can "simulate intelligence" well enough under any conceivable circumstance (and yes, all actual human beans will fail here), how can you say that it cannot "actually think"?
While the limits in this case were admittedly rather strong, this does not form a fundamental objection to the test. The bar can be set progressively higher.
The problem with positively defining intelligence is that all such definitions seem to end up begging the question and are therefore unsatisfactory to somebody. Which is what makes the Turing test philosophically interesting.
The test is an adversarial game, and both sides can attempt new strategies of fooling or seeing through the opponent. I think your claim about being able to pass any set of limits is rather bold, and I would be curious about arguments for it.
We're already in the territory where there are certain applications for these chatbots, and there is no fundamental argument why we this line of research could not progress further until it reaches its goals. I don't mean to say I believe we'll have AI soon, just that you can't fault the Turing test for not doing what it should.
What do you mean when you say there is a problem with positively defining intelligence? Do you think we should define it negatively? There is a lot of criticism of the Turing test, but alternative proposals are exceedingly rare.
http://gizmodo.com/this-is-the-first-computer-in-history-to-...
That the Turing Test is still used is proof that we still don't understand how to even define sapience. Without a definition and concrete, testable qualities, how can we possibly hope to ever build artificial sapience. As a result, we continue to see these toys that are little more than parlor tricks.
Any true test should include looking behind the curtain. "I know you're artificial--I can see the processes working--yet I have doubts that what I'm seeing is real."
In other words, real success is the tester believing he is being fooled when he is not, rather than fooling the tester into believing it is real.
It made me think that we need to divide the turing test for different people: by age, by need, etc.
Generally it involves using "average people" [...] it
should consist of computer science experts instead
I don't agree. The prominent reason for "dumbing down" the judges follows the same reasoning behind decisions made regarding what constitutes "adequate" encryption. How can we gauge what will honestly happen out in the real world, today?Consider DES. It was deemed inadequate, but how to prove it? The EFF came up with a budget estimate based on what was reasonably affordable for a group of attackers, and built a machine capable of cracking DES within those budget constraints.
http://w2.eff.org/Privacy/Crypto/Crypto_misc/DESCracker/HTML...
So, let's apply similar reasoning to the concept of AI. If a group of people were to build an AI, and use it as an adversary against ordinary people, how difficult is it to manipulate and deceive ordinary people into taking action, and is it feasible to do so with AI?
Do I believe that it's intelligent? Well, I wouldn't confer human rights to it. It doesn't carry the weight of emotional investment that a domestic pet might.
But that's the sort of thing to stay wary of. Average people becoming emotionally invested in silly things. People being tricked into carrying around an urn full of ashes, and believing that a convincing AI truly represents their late relatives, and similar sorts of tomfoolery.
The problem with "looking behind the curtain" is that it traditionally boils down to what I like to think of as the subtle fluid model of intelligence. If you know what is behind the curtain, then obviously it can't be intelligent, because it is not running on the right hardware, for gooey definitions of hardware, or because it doesn't have some Homunculus of Definite Understanding (Hi, Serle!), or because we can see behind the curtain and know what it is doing. Obviously, if we know what it is doing, it is not intelligent, right?
Seems that for many people though, especially in CS, intelligence gonna be something that works in a way that is complex enough for them to not understand how it works.
The main point I took away is that he feels that consciences is an ordinary biological process, and a simulation of that process is not the same as the process. In the same way that a computer simulation of a stomach digesting food isn't the same as an actually stomach digesting food. No matter how good the simulation is, it doesn't actually digest food. So a simulation of consciences, isn't actually consciences, and doesn't have a personal subjective experience.
That criterion implies that once we have a much better understanding of the brain that humans might start failing it. What if you have a neuronal tracedump for a person involved in a conversation?
I asked him a few questions like where he lived, his name, if he has brothers or sisters, if he wears glasses, etc. Eventually he started asking me questions like what I did for a living and where I lived. He also managed to form questions based on my answers.
It almost felt like a conversation. I can honestly say, I've never thought that before while talking to an AI. So far I am pretty impressed.
I'm going to take this with a hug pinch of salt until I've read the transcripts, due to the involvement of famous publicity-hound Kevin Warwick.
I'm pretty sure you're human, with a high degree of confidence, much higher than 50%.
So to are you more human than human?
But in this variant, "Do you believe the entity you spoke with to be a human or a program?", a 100% threshold is theoretically achievable. That means you're entirely correct: a 51% vote of confidence is decidedly less human than human. Additionally the 30% threshold they've used is laughably low in this context. Without a control group, even an 100%-confidence outcome probably says more about the beliefs of the judges than the ability of the program to simulate a human.
I don't mean to take away from the no-doubt impressive achievement by the team behind Eugene. I just take issue with the hyperbole in its reporting. But, ya know, the media will be the media, and academics gotta get research grants.
Transcripts would be handy. I doubt a conversation with a 13 year old boy is a good way to measure AI? It's not the best metric to have but it is the most universal and most widely agreed on that we have. It seems like we are happier with small gains in mimicry for now, since real intelligence is hard. Really hard.
It's important to note the true meaning behind this school of thought - that mimicry and "true" intelligence are actually equivalent. Behaviorists believe this, and the validity of the Turing Test along with it.
"Asking whether a computer can think is about as useful as asking whether a submarine can swim."
I think this is fairly insightful and relevant.
I digress. Most of the times, mimicking Nature works very well. We tend to always remember our most spectacular failure, but that's a selection bias.
You don't have to go very far looking for examples. The fictional Nautilus submarine was named after the animal Julio Verne copied its depth control system (the same one all real submarines use till today). Also, the first submarines' shape copied whales, a design that was adapted (but not completely replaced) because it does not work as well witout also copying their propulsion system.
By the way, birds wings are better than planes ones in several important ways (but worse in a few others). We don't copy them because we don't have good enough materials, not because it's not a good idea.
This makes me think that AI research isn't going to be a gradual process of research being built on other research but it will be a eureka or an ah-ha moment that changes everything.
A researcher claiming to have passed the Turing test instantly discredits himself as a prestidigitator looking for PR buzz. The present article is a textbook example of this.
As a side note, if you are focusing on disembodied, language-based human-like intelligence, then the paradigm you operate in is many decades behind. The Turing test was conceived at a time when the notion of thinking machines had just started to emerge --a very different time from today, where we have 60 years of AI research behind us. The Turing test has been irrelevant for longer than most AI researchers have been alive. I have never seen it used for anything else than smoke-and-mirrors PR operations.
There's an element of behaviorism which stretches back to Descartes - how can you know that I exist? That I think? You can only observe through my behavior; that my behavior mimics yours.
How then can we judge machines any other way?
Here's a fun idea for several intelligence tests we can call the VLM tests that have nothing to do with two way conversation like the Turing test.
Given "a machine intelligence" spin up a couple million of them in a "fun" simulation environment and see how much thermodynamic dis-equilibrium they generate by whatever social interaction they see fit to apply to each other. Is it as interesting (aka thermodynamic dis-equilibrium) as a GoL or a real world anthill or a Dwarf Fortress or a Google Earth? Is their simulated culture as interesting to read as HN, or as dumb as youtube comments (which used to be the gold standard of dumbness in social media)
Assuming you can crack the literary code (if any exists) another game to play is extract the meme-flow of a culture of AI vs a culture of 4chan and vote for whichever meme came from a more intelligent group. This is Turing-ish WRT human observers majority vote and such, but is completely non-interactive, merely humans, or even trained sociologists, trying to figure out given two memes which is more intelligent.
Getting out the Sherlock Holmes hat, its possible to determine if an artifact came from an intelligence without talking to the intelligence for awhile. I suspect archeologists have really fun debates on this topic. Is this a stone hammer or merely a peculiar river rock, etc.
A better "Turing test" proposed a few years ago was, "develop a team of robots that play soccer so well they can win the World Cup". While not the best possible driver of AI research, the quest for such a goal would still drive AI research orders or magnitude better than "trick random people to mistake a chatbot for a human".
However, in this case, "Eugene" claimed English as his second language, which seems as close to cheating as it gets.
How is this any better that "build a bot that can beat Chess Masters" ?? Well, apart from the part where this includes multiple agents, and swarm robotics, etc. But frankly, setting a game as a bar actually covers relatively less scope in my opinion. Conversational bot actually does cover a LOT of scope if you think about all the possible ways the conversation can be taken. In fact, we as humans use conversation to judge other humans' intelligence as well. E.g. job interviews.
Though perhaps an advanced turing test could include performance in a variety of social situations like...
- convince employers to hire you, - Convince a girl to go out with you - Convince a customer to buy something - Debates... (Presidential debates by AI would be interesting)
One of the ironies of AI is that trivial tasks for humans like recognizing a soccer ball, moving around without falling over, or, yes, holding a conversation are very, very difficult while difficult tasks, like playing chess or finding the best route on a map, are relatively trivial.
But look at me, ascribing thoughts and intentions on a black box that I cannot be sure really exists.
Judge: Did you hear about the Irishman who found a magic lamp? When he rubbed it a genie appeared and granted him three wishes. “I’ll have a pint of Guiness!” the Irishman replied and immediately it appeared. The Irishman eagerly set to sipping and then gulping, but the level of Guiness in the glass was always magically restored. After a while the genie became impatient. “Well, what about your second wish?” he asked. Replied the Irishman between gulps, “Oh well, I guess I’ll have another one of these.”
CHINESE ROOM: Very funny. No, I hadn’t heard it– but you know I find ethnic jokes in bad taste. I laughed in spite of myself, but really, I think you should find other topics for us to discuss.
J: Fair enough but I told you the joke because I want you to explain it to me.
CR: Boring! You should never explain jokes.
J: Nevertheless, this is my test question. Can you explain to me how and why the joke “works”?
CR: If you insist. You see, it depends on the assumption that the magically refilling glass will go on refilling forever, so the Irishman has all the stout he can ever drink. So he hardly has a reason for wanting a duplicate but he is so stupid (that’s the part I object to) or so besotted by the alcohol that he doesn’t recognize this, and so, unthinkingly endorsing his delight with his first wish come true, he asks for seconds. These background assumptions aren’t true, of course, but just part of the ambient lore of joke-telling, in which we suspend our disbelief in magic and so forth. By the way we could imagine a somewhat labored continuation in which the Irishman turned out to be “right” in his second wish after all, perhaps he’s planning to throw a big party and one glass won’t refill fast enough to satisfy all his thirsty guests (and it’s no use saving it up in advance– we all know how stale stout loses its taste). We tend not to think of such complications which is part of the explanation of why jokes work. Is that enough?
Dennett: "The fact is that any program that could actually hold up its end in the conversation depicted would have to be an extraordinary supple, sophisticated, and multilayered system, brimming with “world knowledge” and meta-knowledge and meta-meta-knowledge about its own responses, the likely responses of its interlocutor, and much, much more…. Maybe the billions of actions of all those highly structured parts produce genuine understanding in the system after all."
I'm sure they didn't get anywhere near this with their 13-yr-old simulation. But this gives an idea of the heights AI has to scale before it can regularly pass the Turing Test.
Though your right, and if a computer were to try to imitate a human, a better strategy would be about as slovenly as my post is.
[edit: Ah, I had the details wrong, see https://en.wikipedia.org/wiki/Turing_test ]
If no one had heard of the rules in advance, rugby would be a pretty decent test of general physical prowess. Maybe not 100% perfect, but out of a population of 1000, the 50 best rugby players would probably match most people's top 50 list of physical specimens well enough.
But, once you have people training and optimizing for it, you find that (a) training for rugby specifically matters. (b) Rugby is optimizing for a particular set of physical characteristics.
Chat bots designed to win the game are basically designed to fool people into thinking that they're human because that's the game. It isn't really a good proxy for consciousness.
re: Chat bots designed to win games: Some say that's exactly what we are! - The Social Brain Hypothesis of the evolution of human intelligence suggests that the reason our brains grew so big was that intelligence (via ability to deal with social groups) became a large factor in reproductive success.
http://en.wikipedia.org/wiki/Evolution_of_human_intelligence...
I think the focus on Turing tests is interesting and has definitely expanded knowledge in this area. But, it is now an area within the search for artificial consciousness. It no longer works as a test for it as it would if a computer just happened to stumble on the test and pass it.
That said, I do thing that where we are visa a vis the Test is a cool benchamark. I would be over the moon if one of the Turing bots got to the point where it could do a job, like being a customer support bot.
Hopefully someplace slightly north of the Turing test goal post there will be commercial goal posts to encourage development, hopefully a conversational user interface. A convincing chatbot as a user interface would present lot of very interesting challenges.
1) Consciousness is really hard to define so the Turing Test is a handy workable yardstick that AI can use as a milestone until we get a proper working definition of consciousness
2) (Hard-AI, behaviourist position) Appearing to be conscious and being conscious are the same thing. Hence the Turing Test is about as good a definition of consciousness as we are ever likely to get. Perhaps it could be tightened up a little - insisting on really long conversations with lots of complexity etc. But a good judge running a test over a longish time period would see to that.
If it came up with any analysis of the joke that was vaguely correct, it's doing much better than anything out there. If it came up with any analysis of the joke at all, I'd be surprised.
That isn't to say all humans have to have meta-knowledge, but the test passing would be more convincing if the AI could do something most everyone can do, like explain a joke.
If a decade or so of social media (whatever that means) has proven anything, its that very little intelligence occurs in virtually all conversations.
The meta Turing test is being failed by many people who think it (a concrete implementation of it) means something. Much like actually building a well sealed box with a cat, a radioisotope source, and a geiger counter wouldn't actually be a "great step forward for Quantum Physics" in 2014. Any more than making a little anthropomorphic horned robot and having him divert fast "hot" molecules one direction or slow "cold" molecules another would be a great step forward for thermodynamics in 2014.
The value of a thought experiment is realized when its proposed, not when someone makes a science fair demonstration of the abstract idea.
Surely it depends on who the human judges are. It seems a bit unfair that the judges normally have IQ > 100 and the other humans have IQ > 100.
I strongly suspect that some simplistic AI (alicebots, for example) would beat the Turing test if the human judges had IQ between 90 and 105. (Especially if we're using the limited 30% rule above).
Getting bots running on some Facebook groups might be interesting.
In short they did nothing.
There is actually more than one bot, which has been claimed to have passed the Turing Test before. Cleverbot is one of them [1]. There are also several competitions, but I believe the most reputable and long standing one is the Loebner Prize [2]. The bot that currently holds the Loeber Prize is Mitsuku [3].
Anyway, you can chat with Eugene at [4], I gave it a try. I believe there is one thing that the creators of Eugene got right. When chatting with other chat bots, I usually in a situation where the bot says something, I ask I followup question (like "Why?"), and it gives a generic answer like "Because I say so" or "I don't know". Eugene does the same but will ask a unrelated followup question right together with the response. That way at least there is not a weird pause in the conversation.
[0] http://en.wikipedia.org/wiki/ELIZA
[1] http://www.geekosystem.com/cleverbot-passes-turing-test/
There are lots of other domains where I would be entirely happy to know that I was talking to an AI, if the answers I was getting were significantly better than most human experts in that domain.
SO the result can very depending on different conditions. :) Highly non deterministic
This is not passing the Turing test by any stretch of the imagination.
Surely the ability to trick a human into believing an AI is a human is a milestone, but it was with an AI specifically optimized for this task. The deeper question is if the passing of the Turing Test in this case means we should ascribe consciousness to the bot, and I think none of us are willing to affirm it yet. I would suggest that this discrepancy is caused by the “measure becoming a target” and losing its ability to be a “good measure.” I guess this is why there is such a critical distinction between Artificial Intelligence and Artificial General Intelligence, which is where the Turing Test would have more weight.
(And this is extended from another form of the imitation game, where the goal is to imitate being male, where participants are male and female)
Have anyone been able to find any more concrete information (and perhaps some transcripts)? If not I hope someone will set up a new test, and invite "Eugene" to participate.
[1] http://arstechnica.com/information-technology/2014/06/eugene...
[edit: We may be given some hints from the wikpedia article on the turing test: https://en.wikipedia.org/wiki/Turing_test#Imitation_Game_vs....
"Huma Shah and Kevin Warwick, who organised the 2008 Loebner Prize at Reading University which staged simultaneous comparison tests (one judge-two hidden interlocutors), showed that knowing/not knowing did not make a significant difference in some judges' determination. Judges were not explicitly told about the nature of the pairs of hidden interlocutors they would interrogate. Judges were able to distinguish human from machine, including when they were faced with control pairs of two humans and two machines embedded among the machine-human set ups. Spelling errors gave away the hidden-humans; machines were identified by 'speed of response' and lengthier utterances." ]
http://gizmodo.com/5921698/what-its-like-to-judge-the-turing...
(linked from http://gizmodo.com/this-is-the-first-computer-in-history-to-... )
I would also add, that the real prove must include a topic that the artificial person was not programmed for. (not like a Bayesian filter that "develops" by "learning" new facts about a fixed topic).
Learning, developing, evolving, that are the real marks of living and of intelligence (since, I would not part between intelligence and living).
Speaking is only a way of getting into the stage. Once into the highlights you must prove you are a leader or, if you decide so, that you are able to gain the attention of your audience to emphasize something important that previously was not perceived as such. That is speaking is an art, is not about explaining a plot but about creating a story.
Move on people, it's just a cheap PR stunt.
But people will continue to dismiss the state of the art and deny that computers have "real" intelligence, the same way they did when the computer defeated Kasparov, the same way they did when we saw Googles self driving cars, the same way they did when a computer won on Jeopardy, and now with the Turing Test. Even when we have robots that look and act exactly like humans, many people will say that they are not "really" intelligent and dismiss the accomplishment. They will still be saying that when AIs twice as smart as people arrive and they have to figure out what to do with billions of what will then be, relatively speaking, mentally challenged people.
So how long until the creator can pass the Turing Test?
(Normally I'd ignore that, but given the subject matter...)
Some forms of Turing test are trivially passable with dumb enough humans.
Edit: Oh, and considering how many fluff press articles are showing up about this [1], the submitter also showed exemplary taste in source selection. Yay, stevejalim!
1. https://hn.algolia.com/?q=turing+test#!/story/sort_by_date/0...