Edit: this tool is as reliable as a magic 8-ball
When a human is miscategorized as a bot, they could find themselves in front of academic fraud boards, skipped over by recruiters, placed in the spam folder, etc.
To make it more concrete on work I am very familiar with: breast cancer screening. If you had a model that outperformed human radiologists at predicting whether there is pathology confirmed cancer within 1 year, but the accuracy was not 100%, would you want to use that model or not?
I agree that there are places where we shouldn't put AI and that checking whether something is an LLM or not is one of them. However I think the sentence above takes it way too far and breast cancer screening is a pretty clear example of somewhere we should accept AI even if it can sometimes make mistakes.
> When a human is miscategorized as a bot, they could find themselves in front of academic fraud boards, skipped over by recruiters, placed in the spam folder, etc.
Is the problem here the algorithms or how people choose to use them?
There’s a big difference between treating the results of an AI algorithm as infallible, and treating it as just one piece of probabilistic evidence, to be combined with others, to produce a probabilistic conclusion.
“AI detector says AI wrote student’s essay, therefore it must be true, so let’s fail/expel/etc them” vs “AI detector says AI wrote student’s essay, plus I have other independent reasons to suspect that, so I’m going to investigate the matter further”
Two people can buy the same product yet use it in very different ways: some educators take the output of anti-cheating software with a grain of salt, others treat it as infallible gospel.
Neither approach is determined by the product design in itself, rather by the broader business context (sales, marketing, education, training, implementation), and even factors entirely external to the vendor (differences in professional culture among educational institutions/systems).
It can be used for some decision (i.e. not critical ones), but it should NOT be used to accused someone of academic misconduct unless the tool meets a very robust quality standard.
> this tool is as reliable as a magic 8-ball
Citation needed
Nearly everything doesn't give 100% accurate results. Even CPUs have had bugs their calculation. You have to use a suitable tool for a suitable job with the correct context while understanding it's limitation to apply it correctly. Now that is proper engineering. You're partially correctly but you're overstating:
> A tool that gives incorrect and inconsistent results shouldn’t have any part of a decision making process.
That's totally wrong and an overstated position.
A better position is that some tools have such a low accuracy rate that they shouldn't be used for their intended purpose. Now that position I agree with it. I accept that CPUs may give incorrect results due to a cosmic ray event, but I wouldn't accept a CPU that gives the wrong result for 1/100 instructions.
Meanwhile, the leading commercial tools for plagiarism detection often flag properly cited/annotated quotes from sources in your text as plagiarism.
On the other hand, an opaque LLM detector that just prints “that was from an LLM, methinks” (and not e.g. a prompt and a seed that makes ChatGPT print its input) essentially cannot be proven false by an author who hasn’t taken special precautions against being falsely accused, so the bar for sanctioning people based on its output must be much higher (infinitely so as far as I am concerned).
The whole silly concept of an "AI detector" is a subset of an even sillier one: the notion that human creative output is somehow unique and inimitable.
The AI detection tool fails both as it has a low accuracy and could ruin someones reputation and livelihood. If a tool like this helped you pick out what color socks you're wearing, then it's just as good as asking a magic 8-ball if you should wear the green socks.
If you are asking, is this LLM text Human generated, and it says Human (yes), then it is false positive.
If you are asking is this LLM generated text LLM generated, and is says and it says Human (no), then it is a false negative.
So I think the only thing a mythical detector could determine would be LLM, or non-LLM, and let us take it from there. But detectors are bunk; I've had first-hand experience with that.
Sure, it's theoretically possible to add two noisy signals that are uncorrelated and get noise reduction, but is it probable this would be such a case?
It all depends on the properties of the signal and the noise. In photography you can combine multiple noisy images to increase the signal to noise ratio. This works because the signal increases O(N) with the number of images but the noise only increases O(sqrt(N)). The result is that while both signal and noise are increasing, the signal is increasing faster.
I have no idea if this idea could be used for AI detection, but it is possible to combine 2 noisy signals and get better SNR.
You can't detect "truth" from that, but you can often tell (i.e. with better accuracy than chance) whether or not a subject is able to give a confident, uncomplicated yes-or-no to a straightforward question in a situation where they don't have to be particularly nervous (which is why it's not very useful for interrogating a stressed criminal suspect, and should absolutely be inadmissible in court).
But everyone knows that it's not very reliable in almost every circumstance it's used. My point is that while only marginally better than chance, it's still better than chance, unlike the OpenAI's detector, which is significant worse than chance.
You can detect indicators of stress... or hot weather... or stage-fright (admittedly a form of stress)... or too much caffeine... or an underlying (maybe undiagnosed) medical condition, etc. So it does not even necessarily measure "stress".
It's about as useful as the so called "fruit machine" which they used to test for homosexuality[0], in that it is utterly useless while at the same time can be quite ruinous for people. People have been fired over polygraph "fails", and while not admissible in courts, people probably have been fingered for crimes after they failed polygraphs. Also, criminals have gone free after passing polygraphs[1].
>But everyone knows that it's not very reliable in almost every circumstance it's used.
You and I may know that. But a lot of people actually do not. That's why it's still used. Either because people administering those tests think it's "good science", or because those people administering it know that while it's all bullshit the person they are testing might not know that and break down and admit to things. Remember that fake polygraph on the show The Wire, which was just a copier they strapped to the suspect. If I remember correctly that was based upon true events.
A quick google shows e.g. you can hire "polygraphers" to e.g. "test" if your partner was unfaithful, making claims such as: "However, assuming that you have a good polygrapher with a fair amount of experience in working with betrayal trauma, you're going to get results that are at least 90% accurate or better."[2]
The US (and probably a lot of other) government(s) like their polygraphs very much, too[3].
> you can often tell (i.e. with better accuracy than chance) whether or not a subject is able to give a confident, uncomplicated yes-or-no to a straightforward question in a situation where they don't have to be particularly nervous
Uhmm, if somebody sat me down in a room, strapped all kinds of "science" to my body and then asked me questions, I'd be quite nervous regardless of whether I am truthful or not. In fact, I'd be even more nervous knowing it's a polygraph and bullshit, because I cannot know if the person administrating it would know that too.
If that somebody then asked me "Have you ever killed a prostitute?", or "Have you ever colluded with the enemy?", or "Have you ever cheated on your partner?", or "Have you ever stolen from your employer?", for example, my stress would certainly peak despite being able to confidently and truthfully answer "No!" to all of those questions. And I am sure the polygraph would "measure" my "stress".
[0] Yes, that was a real thing too. https://en.wikipedia.org/wiki/Fruit_machine_(homosexuality_t...
[1] E.g. the Green River Killer Gary Ridgway passed a polygraph, so the police turned their resources to another suspect who failed the polygraph. That was in 1984. Ridgway remained free until his arrest in 2001. He killed at least 4 more times after the investigation stopped focusing on him after that "passed" polygraph.
[2] https://www.affairrecovery.com/newsletter/founder/use-abuse-...
[3] https://support.clearancejobs.com/t/the-differences-between-...