Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
I tried a chapter just now and got human doing that, but I'm not invested enough to run a hundred samples today. But it sounds like it would be an alright way to audit it? I will confess I'm pretty skeptical you could ever eliminate false positives here though. I can often get an ai sense from some writing on my own but I doubt it would be better than 90% accurate, and "ai plus human editing" might screw with that anyway, stuff like that. I would have preferred we just never developed this kind of thing so I wouldn't have to guess.
if they used older texts as training data, to some extent pangram would just be an age classifier for writing style.
Also, this doesn't even consider the case where people use LLMs to translate their original works. Or people that use it for spelling/grammar checks.
Personally, I believe these checkers do more harm than good. Any false positive can ruin someones life.
I recently heard someone say "that's genuinely the exact solution I was looking for" and had to do a double take.
https://www.pangram.com/research/model-card/pangram-4
> Pangram 4 achieves a 0.0041% false positive rate (roughly 1 in 24,000) on 1,000,000 human-written English FineWeb evaluation examples
> Overall False Negative Rate is 0.3396% on English AI generations (26 generator models)