The answer: "One way to make it clear that a text is not generated by AI is to include a joke or riddle that is difficult for a machine to understand or generate. For example, you could include a joke that relies on wordplay or a complex pun. Here is an example:
Why was the math book sad?
Because it had too many problems."
Good thing ChatGPT doesn't know any complex puns or witty wordplay!
The thing is, I'd see these answers as similar to everything else the program produces. A bunch of claims from the net cobbled together - I've read a number of Sci-fi novels and stories where "inability to understand humor" is the distinguishing quality of an AI (I'm guessing it extrapolated "hard create" from "hard to understand"). But that doesn't seem to be playing here where the AI mostly running together things humans previously wrote (and so it will an average amount humor in circumstances calling for it).
A reasonable answer is that the AI's output tends to involve this running-together of common rhetorical devices along with false and/or contradictory claims within them.
-- That said, the machine indeed did fail at humor thing time.
The question here is this an actual AI only failure mode. Are we detecting AI, or just bullshittery?
I would say that human knowledge involves a lot of the immediate structure of language but also a larger outline structure as well as a relation to physical reality. Training on just a huge language corpus thus only gets partial understanding of the world. Notably, while the various GPTs have progressed in fluency, I don't think they've become more accurate (somewhere I even saw a claim they say more false thing now but regardless, you can observe them constantly saying false things).
And the idea that the computer would “try” to come up with an example that would trick a computer is itself a little funny, in that it has fallen into giving itself a preposterous task.
But it did definitely fail at clever wordplay.
There sure is some obscure discussion forum where users talked about that or some amateur writer that published online something in those lines. ChatGPT is just a statistical device selecting randomly from previous answers.
Of course, in real time the attempts at humor often fall flat and might give away flawed thought processes, although I personally have found them to be often insightful, (containing a seed of humor) even when they're not funny. It could be a useful technique when actually having a conversation, a form of Voight-Kampff test, but I don't think it will do anything to let you know if the content was generated by AI and then just cherry picked by a human.
In two years? Look where the space was two years ago. I think many things will have to change.
I can just fine-tune a large scale model on a small downstream task, or use creative choices of decoding settings (high temperature, alternative decoders like contrastive/typicality sampling), to fool the existing methods.
Some times, though, the answers are false but plausible.
* canned laughter *