The Loebner prize completely misrepresents Turing’s paper
yaxu.org
yaxu.org
I propose to consider the question, "Can machines think?" This should begin with definitions of the meaning of the terms "machine" and "think." The definitions might be framed so as to reflect so far as possible the normal use of the words, but this attitude is dangerous, If the meaning of the words "machine" and "think" are to be found by examining how they are commonly used it is difficult to escape the conclusion that the meaning and the answer to the question, "Can machines think?" is to be sought in a statistical survey such as a Gallup poll. But this is absurd. Instead of attempting such a definition I shall replace the question by another, which is closely related to it and is expressed in relatively unambiguous words.
So what he's really saying is that the question "Can machines think?" is too ambiguous since the words "machine" and "think" aren't well-enough defined. We could use a Gallup poll to find out what people think the words mean, and then try to answer the question based on those definitions, but that would be absurd.
What's really "wrong" with the Loebner Prize is that it's a very simplified Turing test (five minutes of conversation) which doesn't promote the development of actual thinking machines any more than a paper airplane competition promotes the development of transcontinental airliners. Eventually somebody will write a chatterbot which is capable of fooling most of the judges into thinking it's human for the duration of a short conversation (if you've read the transcripts they rarely go on for more than a dozen lines or so) but it will be a bunch of clever tricks rather than an AI, and a serious attempt to figure out whether it's intelligent (e.g. by telling it a story and then asking it basic reading-comprehension questions about the story) will quickly show that it isn't.
Really?
If all we had to begin with were paper airplanes, then a paper airplane competition would certainly promote transcontinental airliners. Before powered flight, studying paper airplanes was certainly one valid way to make progress. And the limits of today's AI in many ways make it seem pretty close to paper airplanes - so seeing if they can every fly across a room seems entirely valid.
SHEESH, it's hard to find even a metaphor in which this controversial prize doesn't advance our knowledge.
Consider, human behavior is very ad-hoc. AI as a field has neglected ad-hoc interaction, preferring more logically specified activities. If a computer could hold an effective five minute chat, it would be so far above what currently exists as to be breath-taking. If that's done with 'tricks', it's time to start understanding the 'tricks' rather going on and on with the systematic or whatever approaches we might have thought were the proper way to do this.
Oh, and the reading test sounds easier than the interaction test - MUCH easier.
Incidentally, I just had a chat with last year's winner elbot (see elbot.com) and it's got some pretty clever special-case interactions built in. For instance:
me: What is 4 times 4?
elbot: The answer is: 17!
me: Haha, very clever.
elbot: I'm glad you agree, most people think the answer is 16.
a) Reverse-engineering a human brain at the neural level, or b) Some kind of evolution
whereas I'm reasonably sure that the Loebner Prize could, given enough money, brains and time, be won by a large bag of special-case tricks based on careful observation of how the judges tend to behave in practice.
Each program, upon startup, must provide a “browse”
function to select a directory. Communications shall
be by means of the creation, detection, and deletion of
sub-directories within the specified communications
directory.
To simulate a key press the entry program must create a
sub-directory within the communications directory with
the following format:
“time.keypress-name.extension”
where time is a monotonically increasing 18 digit number
(in lexical and numerical order) (i.e. zero filled to the
left) to be retrieved from the system clock and expressed
as milliseconds past some initial time as defined by the
system clock.
“keypress-name” is either a single letter (case sensitive)
or the name of the special character, as appended to
these rules.
The extension is “.other”
For example: “000001234567890123.bracketleft.other”
To detect a key press by the judge, the program must
detect, within the communications directory a
sub-directory with the same format, but extension
“.judge” and then must remove or delete the judge’s
sub-directory from the communications directory.
It's not the end of the world, but at the same time it's a fairly nasty way of doing something that should be relatively easy. I entered in 2002 before they introduced this and have been hesitant to enter again in recent years partly because I don't want to have to deal with the hackery of the 'communications protocol'.Chris McKinstry (the guy behind the MindPixel project) got fed up with them and made the Minimum Intelligent Signal Test (MIST), which is more objective than the original Turing test and may be the simplest way to test for general, common-sense intelligence. http://en.wikipedia.org/wiki/Minimum_Intelligent_Signal_Test
Edit: forgot link. d'oh.
http://en.wikipedia.org/wiki/Minimum_Intelligent_Signal_Test
It's not as if the machines are acing the Loebner version of the test or anything. Obviously, the test is dumbed-down to allow Bronze metal winners. Seems like a logical even if so much else about the contest seems mired in the persona of Hugh Loebner.
Turing's intent was to define a sufficient condition for labeling behavior "intelligent", and I think a high score on Loebner's setup with a competent interrogator does this better than a pre-defined set of MIST questions.