My experience in testing actual AI written content on willing participants is that people are entirely useless at detecting AI written content with any reliability whatsoever.
And I believe my experience is something expected. People are also certain kind of a neural network. If an LLM system is trainable to be a decent detector, I don't see a reason why at least some people couldn't be.
Similar as with coding, yes, halting problem!, but we've been always reviewing code nonetheless.