[1] https://cs.nyu.edu/faculty/davise/papers/WinogradSchemas/WS....
[1] https://cs.nyu.edu/faculty/davise/papers/WinogradSchemas/WS....
[Session starts]
<Challenger> Hi.
<Judge> Hi. Now, listen to me. My job here is to judge whether you are a human or an AI. I also know that your job is to try and persuade me to believe you are a human, no matter which one you really are. I happen to have a few questions that are easy for humans and hard for AIs. So if you are a human, answering them shouldn't be a problem to you, but if you avoid answering, I'm going to judge that you are an AI, no excuses. Here is the first question: ...
[A pattern of Winograd schema challenges ensues.]
So as we can see, the Turing test doesn't have to be any random chitchat nor it doesn't have to pretend to be "normal human-to-human-like discussion", as the Judge has the ability to participate and steer the conversation and zero in on the areas that are most likely to expose the AI.
However, if the human participants don't have any penalty from being labeled as AI, this doesn't obviously work, as they don't have any pressure to prove anything. So this kind of an embedding wouldn't work with "imitation game" -like rules, but it would work with "prove your human-ness" -like rules.
So, you are limited to questions that most people will answer correctly. Further, if you find some unusual question that works today someone can just add it to the program for next time.
The trophy doesn't fit into the brown suitcase because it's too large. What is too large? A: The trophy B: The suitcase
If you keep writing specific solutions to this class of questions it's eventually had to come up with a simple question that's still unknown.
Further you can't get around that by making ever more complex questions because humans will eventually start messing them up.
And if your version of the test is "heads it's an AI, tails it's human" then any AI's that are classified as human will have "passed the Turing test."
The original test specifically had exactly one human and one AI. So, if the judge is forced to do a coin flip that really is success. If the judge does a coin flip because they are lazy then that's not a Turning test.
That's why its useful.