Frequent users of ChatGPT are robust detectors of AI text
arxiv.org
arxiv.org
The explanation for the difference is that automated discrimination has relied mainly on structural factors such as average sentence/paragraph length and frequency of stock words/phrases and certain parts of speech. Human evaluators look at content factors such as repetition of ideas, less precise wording, generalizations rather than concrete examples, overall conceptual coherence, and factual errors.
There is some discussion of LLMs and models in another thread. (https://news.ycombinator.com/item?id=44625629)