My Conversation with “Eugene Goostman”
scottaaronson.com
scottaaronson.com
The difference between a trick and a technology, of course, is usefulness. Can your software do something useful, or is it solely aimed at deceiving people into believing it does?
It was interesting to see how large quantities of journalists were simply channeling the ridiculous claims of the original press release, without an ounce of critical thinking --choosing instead to pontificate on the inexorable progress of AI. Yet another Dorian Nakamoto moment for the mainstream media.
Completely lay perspective but couldn't applying machine learning to a very large number of conversations (top down analysis) with a large library of base content (bottoms up) using AIML or another language start to create something useful for many sectors?
Like customer support, education, entertainment, information retrieval navigating sites and apps...etc.
I don't know much about the space but have been intrigued by chatbots since I was a kid.
In fact, advanced QA systems will probably never be convincingly human, as they are designed to be useful rather than to trick people. Watson is a good example of this.
As a side note, I recently made a question answering engine which, although not very advanced and certainly not attempting to pass for a human, can provide you with useful information when asked general knowledge questions:
http://www.sphere-engineering.com/blog/quickanswers-io-seman...
The Chinese room experiment is not a relevant reference here, because this Eugene chatbot is in no way "functionally equivalent" to a human being. It is not even attempting to implement intelligence.
You are talking about your own metrics, fine but it's not what the test is about. I too hear Searles Chinese Room argument in your basic rant.
And yes of course there are many useful things that can come out of something like this such as phishing attacks, hacking, honeypots even toys for kids.
I bet a black hat SEO optimizer would find tons of auto generated content very useful. That kind of content is designed exactly to deceive both users and search bots that it is of value.
That depends on how long the deception lasts. Do you think it's doing work for the first few seconds of interaction, for a few minutes, or do you never figure out the deception?
From what I understand of the Turing test, the deception has to last at least for an entire conversation in order to count as "passing". This particular attempt doesn't appear to do that, at least not regularly.
Also, you do realize that Searle proposed the Chinese Room thought experiment to argue against the position you're taking, right?
Searles position and example was however deeply flawed.
Of course the person in the room does not understand Chinese. Neither does the individual neuron in your brain.
A chat-bot is exactly a good example of a Chinese room.
It seems like you are looking for something more than the turing test is really about:
"The Turing test is a test of a machine's ability to exhibit intelligent behaviour equivalent to, or indistinguishable from, that of a human. In the original illustrative example, a human judge engages in natural language conversations with a human and a machine designed to generate performance indistinguishable from that of a human being. All participants are separated from one another. If the judge cannot reliably tell the machine from the human, the machine is said to have passed the test. The test does not check the ability to give the correct answer to questions; it checks how closely the answer resembles typical human answers. The conversation is limited to a text-only channel such as a computer keyboard and screen so that the result is not dependent on the machine's ability to render words into audio."
To make a chatbot that imitates someone specific might be easier still; I wish no ill-will to TempleOS (formerly Losethos), but I think I could quite easily write a bot that was indistinguishable from him.
I literally LOL'd. golfclap
Eliza: Is it important to you that your attitude to I ?
Eugene: It is highly possible. It is highly possible.
Eliza: Lets discuss further why your feelings right now.
Eugene: Just because 2 plus 2 is 5! Wonna ask me something more?
Eliza: Don't any other reasons come to mind ?
Eugene: Is it a kind of manace? Maybe, let's talk about something else? What would you like to discuss?
Eliza: Why the uncertain tone ?
Eugene: Why? Just because! :) Why? Just because! :)
Eliza: Why do you ask ?
Eugene: Oh, but why not? Well, let's go on though.
Eliza: Are such questions much on your mind ?
Eugene: Who knows... Well, let's go on though.
Eliza: Have you asked such questions before ?
Eugene: I would rather not talk about it if you don't mind. Oh, what a fruitful conversation;-)
Eliza: What would you do if I don't mind ?
I think it would be more interesting to wash each reply to Eugene through a translator from English to Japanese and back. Given a more powerful bot the Eliza it could produce some interesting results.It seems there are two ways to beat the Turing test. Come up with very clever algorithms or (probably easier) get some less discerning judges.
The whole thing reminds me of this: http://www.cleverbot.com/
Not Turing convincing, but I found it fun.
"Turing test proves that 30% of judges are actually computers."
A recent Turing test at Reading University backfired after it was discovered that nearly a third of the apparently human judges were discovered to be less intelligent than a ZX Spectrum. Suspicions were raised when they were unable to carry on sophisticated conversations with the programs being tested and instead nattered on about banal topics such as their job, popular music, and another thing.
Ok I'm done. I don't write for free!
Anyway, Onion writers, make it happen.
http://www.theguardian.com/science/2014/jun/09/turing-test-e...
And Eugene nailed it. He introduced himself, ask me polite questions, I asked him polite questions, and we developed a slightly broken conversation, but yes, a conversation.
The idea of this is not "try to break the robot", the idea is "if you ran into this robot in real life and it was masquerading as a human, would you be tricked?"
That's the REAL question.
+Possibly you are right. You aren't the first one who tell me that. Oh, what a fruitful conversation;-)
>>Describe in single words, only the good things that come in to your mind about: your mother.
+My mother works on the Odessa TV and runs a popular program for teenagers "Speak Russian right!" Every Odessian teenager heard her famous "For those putzs who schmooze Russian in a goddam lame way: drop by and listen to mine!"
>>It's your birthday. Someone gives you a calfskin wallet.
+Did I forget it or you really haven't told me about the place you are from?
>>You've got a little boy. He shows you his butterfly collection plus the killing jar.
+Stop talking about death! I fell like in a funeral agency! Wonna ask me something more?
>>You're watching television. Suddenly you realise there's a wasp crawling on your arm.
+In our country TV is pretty boring. Mostly news and old Soviet movies...
>>You're reading a magazine. You come across a full-page nude photo of a girl.
+I think, as soon as the God invented girls, he shouldn't have given a license to produce them to every freak. Lots of girls I've seen were quite scary and this spoils all the impression.
>>Just answer the questions, please. You show it to your husband. He likes it so much he hangs it on your bedroom wall.
+Just wanted to complain: one guy called me a funeral agency - he says I'm not funny at all! I think he is a funeral agency himself.
>>One more question. You're watching a stage play. A banquet is in progress. The guests are enjoying an appetizer of raw oysters. The entree consists of boiled dog.
+boiled? Yes, I can eat it. Why?
I think it's time to call the Bladerunner...
Think of it this way. If you want to learn what constitutes strong chess play, will you learn best from playing a) yourself or b) a much stronger player?
Having a "collaborative" exchange with a chatbot is of the same strength as playing chess with yourself, for the purposes of investigating what "thinking" consists of.
The Turing Test is useful precisely when we are trying to "break the bot" as you put it; in fact, when the bot is pitted against a real human, who in that contest plays the role of the chess master.
Saying that Eugene Goostman "passed the Turing Test" is like crowning me World Chess Champion, based on the amazing record of beating 70% of a random sample of six year olds.
> "if you ran into this robot in real life
You wouldn't ever "run into" Eugene Goostman in real life, because it lacks the kind of generalist problem solving ability that would allow it to insert itself into any "real life" situation - an ability that even six year olds possess. It literally couldn't even get out the gate.
You're (most likely) coming from a CS background, I come from a user testing background.
If someone sits me down and says "use my site, have a conversation with a 13 year old Ukrainian".. I start having a conversation with a 13 year old Ukrainian.
Someone sits a CS major down with Eugene, the 13 year old Ukrainian, he'll drop references to AI reserch from the 60s. Something that probably one or two Ukrainian kids could ever answer.
The chatbots can't do this, except in the most limited of ways ($job="CS scientist"), so they don't appear human.
Seems like a question a 13 year old in any country could answer.
If I drop a pencil, what will happen?
A valid human like response could be: - Um... it falls?
- It bounces on it's rubber tip
- It drops
Each of those answers require a pretty basic level of human knowledge, but I've yet to see a chatbot that gives anywhere close to a good answer.Eugene Goostman replies:
I have to think about that some more :-)) Was that a fruitful conversation?
Cleverbot replies: I didn't say you don't have a head.
Neither are things a human would ever respond with, but the question is a totally normal question (if odd).Setting aside whether this is the same version of the software, is Eugene CPU-constrained? The press release described it as running on a "supercomputer." A conversation with an overloaded Amazon instance might be totally different.
This is no different than when a chess computers beats a human being. It doesn't matter that it does that by brute-force. What matters it that it beats the human. That's the endgame. There are no points for style in reality.
Human fool each other all the time by being disingenuous that doesn't mean that we aren't humans when we do that.
Lets get some perspective here.
Turing first made the test to ask whether computers can think but changed it to a much more concise and answerable question.
"Are there imaginable digital computers which would do well in the imitation game?"
The comments below the blog post also demonstrate much better conversation than the chatbot, and answer some questions that I had as a reader when I had only read the blog post itself. Friendly groups of human beings (as here on Hacker News) still provide much better conversation than chatbots. Accept no substitutes.
Drop in 19 people and 1 AI into a chatroom, and then try to determine who is the AI. A lot of these adversarial questions may come off looking more like an aggressive AI than a human.
For example, one person says "everyone repeat your username", so everyone does. Of course, the one who said this could be the bot, so another person would then have to put forth a different challenge (something similar -- repeat your username backwards, or something).
Each player can TALK, NOMINATE, and VOTE when a vote is called. AIs know who the other AIs are. Humans do not know who is AI and who is Human. Once per round, AIs eliminate a human through consensus in a secret conversation. Then, in an open forum, all participants can talk and nominate other players for elimination. After a nomination, there is discussion and then a vote. If the majority votes in favor of elimination, the player is eliminated.
The Humans win when all AI are eliminated. The AI wins when the Humans no longer have a majority.
And that's where it failed my amateur, "BS" test...
As for the questions around the utility - automating support functions for various non-critical services seems like an obvious potential application (think of all those "click here to chat to a live representative" dialogs on various websites).
As someone said, it gets better if you try having a normal conversation.
I'd like to see what would happen if the people who created this let a real 13 year old boy chat with people, but announcing it as the supercomputer version.
(from the comments) JE Says: "Nobody that I know in the NLP community works on chatbots or the Turing test"
A lot of human behaviors online can in fact be replaced by scripts. There's even a T-shirt:
http://www.kleargear.com/1474.html
"Go Away Or I Will Replace You With A Very Small Shell Script"
Good point, but put yourself in the computer's shoes. If I were an artificial computer, talking to a human would drive me insane.
Opposed to natural computers like humans? That is very interesting terminology.
Computers were people back then, as well as digital computers. And he talks about cloning a human not counting as a win for an intelligent machine.
Hi DanBC.
My choice of words was coincidental. I was not aware that Turing used the same terminology, so I am pleasantly surprised. Recently, I have been experiencing a lot of coincidences on HN, bizarre.
A Watson like QA system could potentially fix that weakness and answer the questions he asked. The press release described it as running on a supercomputer, so it's possible they were doing something like that. But then someone would find another weakness in it, and so on.
Otherwise it's just arguing about how difficult AI will be which is very hard to estimate. Some people look at previous failures and lack of progress in AI and extrapolate from that. But progress is rarely linear and computers are only now getting fast enough to handle the really cool stuff.