[Session starts]
<Challenger> Hi.
<Judge> Hi. Now, listen to me. My job here is to judge whether you are a human or an AI. I also know that your job is to try and persuade me to believe you are a human, no matter which one you really are. I happen to have a few questions that are easy for humans and hard for AIs. So if you are a human, answering them shouldn't be a problem to you, but if you avoid answering, I'm going to judge that you are an AI, no excuses. Here is the first question: ...
[A pattern of Winograd schema challenges ensues.]
So as we can see, the Turing test doesn't have to be any random chitchat nor it doesn't have to pretend to be "normal human-to-human-like discussion", as the Judge has the ability to participate and steer the conversation and zero in on the areas that are most likely to expose the AI.
However, if the human participants don't have any penalty from being labeled as AI, this doesn't obviously work, as they don't have any pressure to prove anything. So this kind of an embedding wouldn't work with "imitation game" -like rules, but it would work with "prove your human-ness" -like rules.