This post dates back to February - the Stanford Institute for Human-Centered Artificial Intelligence (HAI) made their DetectGPT application as a demo and published a paper. They made the ambitious claim that their tool correctly identified human versus authorship in 95% of the cases across five popular open-source large language models (LLMs).
Since then, papers have been published to show the bias against non-native speakers, whose productions were more likely to be classified as AI-generated. Prompt engineers have also found creative ways to "humanise" AI-generated output. Where do we stand today? I can vouch for a significantly lower than 95% reliability threshold - and, by the way, this applies mainly to English - a different detection tool needs to be developed for other languages, or at the very least for other language families.