This has been pretty well studied. Pangram is about as good at an expert human with a 98% detection rate and <2% false positive, for flagging AI generated text.
It's not exactly a wide breadth of varied text for such an all-encompassing task.
You can also just try it yourself I guess really what convinced me was how it perfectly agrees with my own judgement.
Now, I could just assume I'm an outlier within that small false-positive rate. But I find it far more likely that the efficacy of these tools is much lower than advertised and they rely on confirmation bias.