Everything has false positives or negatives. 2 examples doesn't tell anything about how good or bad it is.
For example they have a corpus of older pre-llm text and use that as the example their current model doesn't misclassify human written text. It shouldn't take much thinking to realize why this is a fucking stupid benchmark.
Every day humans use LLMs and read LLM content Panagram becomes more useless because it forces languages to have a stopping point sometime around 2020. If you adopt any LLMism or are one of those unlucky people that already talked like an LLM before LLMs then all your shit is getting marked even though it was created by the human mind and written by human hands.