That'd be if they had a false discovery rate of 1/10,000.
If for instance:
* 100,000 samples are tested
* 100 of which are AI-generated, the rest human-written
* Pangram flags 50 of the AI-generated samples (true positives)
* Pangram also flags 10 human-written samples (false positives)
Then the FPR is 1 in 10,000, but the chance that a flagged sample isn't actually AI (FDR) is 1 in 6.