Immediately saw false positives on content written before ChatGPT. Not difficult to see how the methodology is wrong when it considers restating the thesis in the conclusion to be signal.
Re methodology: Restating the thesis is one signal of many. 77% of the AI posts close that way, but so do 12% of the human ones. The classifier in the paper never decides on one feature, but their combination.