(haven't read the paper)
Doesn't it boil down to a halting-problem kind of argument?
Assume you have a model that can detect AI-generated text. Then you can build a model that emulates the detection model, finds the nearest non-detected text, and outputs it.