Spotting LLMs with Binoculars: Zero-Shot Detection of Machine-Generated Text
arxiv.org
arxiv.org
As an example, one method that I found that works extremely well is to simply rewrite the article section by section with instructions that require to mimic the writing style of an arbitrary block of human written text.
This works a lot better than (as an example) asking to write in a specific style. Like, if I just say something along the lines of "write in a casual style that conveys lightheartedness towards the topic" is not going to work as good as simply saying "rewrite mimicking the style in which the following text block is written X" (where X is an example of a block of human written text).
There are some silly things that will (a) trigger human written text to be detected as a AI and (b) that allow to avoid AI detection, e.g. using broad dictionary tends to trigger AI bots to detect the text as written by AI. So if you are using Grammarly to "improve your writing", then don't be surprised if it gets flagged. The inverse is true too. If you some statistical analyzes to replace less common expressions with more common expressions, AI-text is less likely to be detected as AI.
If someone is interested, I can talk a lot more about hundreds of experiments I've done by now.
If you need a dataset to benchmark against, download any articles from pre 2017. There are a few ready-made datasets floating around the Internet.
I am sure that these algorithms have evolved, but given my past experiments, I sincerely doubt that we are at a point that (a) cannot be easily bypassed if you are targeting them, (b) do not create a lot of false-positives.
As stated in another comment, I personally "gave up" on trying to bypass AI detection [it often negatively impacts output quality], at least for my use case, and focus on creating highest-possible value content.
I know that services like Surfer SEO are continuing to actively invest in bypassing all detectors. But... as a human, I do not enjoy their content and that what matters the most.
> Over a wide range of document types, Binoculars detects over 90% of generated samples from ChatGPT (and other LLMs) at a false positive rate of 0.01%, despite not being trained on any ChatGPT data.
So I'm a researcher in vision generation and haven't read too much about LLM detection but am aware of the error rates you mention. I have questions...
What I'm absolutely surprised by is the use of perplexity for detection. Why would you target perplexity? LMs are minimizing NLL/entropy. Then instruct based models are even more tuning in that direction such that the you're minimizing the cross-entropy as compared to human output (or at least human desired output). Which makes it obvious that it would flag generic or common patterns as AI generated. But I'm just absolutely baffled that this is the main metric being used, and in the case of this paper, the only metric. It also gives a very easy way to fool these detectors since it would suggest just throwing in a random word or spelling mistakes would throw off detection given that such actions clearly increase perplexity. To me this sounds like using a GAN's detector to identify outputs of GANs (the whole training method is about trying to fool the detector!) (Obviously I'm also not buying the zero-shot claim).
If all we're getting is false positives then it can be used to reduce the workload.
If we also get false negatives then we'd be better off using existing techniques (manual or otherwise).
I suppose an issue with this might be that an unknown prompt would add a lot of "hidden" information, but you could probably start from a guess or multiple guesses at the prompt.
It's a clever plan, until the LLMs do some adversarial training....
Perhaps this measurement approximates a human reaction to chatGPT: 'This writing is distinctly indistinct.'
> They did not invent a machine to destroy technology, but they knew how to use one. In Yorkshire, they attacked frames with massive sledgehammers they called “Great Enoch,” after a local blacksmith who had manufactured both the hammers and many of the machines they intended to destroy. “Enoch made them,” they declared, “Enoch shall break them.”
... And another reference for the phrase...
https://www.nigeltyas.co.uk/nigel-tyas-news/post/enoch-the-p...
> And here’s the funny thing. The weapons they reached for to wield and smash the machines were sledge hammers made by ... the Taylor brothers of Marsden. This irony was not lost on the Luddites and as they swung ‘Enoch’s hammers’ to damage his hated machines they cried: “Enoch made them, and Enoch shall break them”.
it's an unwinnable war.
Maybe optimizing for “an amount of imperfections, variability, and subtle but tell-tale clues as to convince an examiner that the writer is human and not a language model.”?
If that’s not enough it could also be asked to read the paper itself and suggest additional countermeasures based on a technical review.
2. How well can it detect if the prompter tries to hide it?
3. How well can it detect if people tend to start writing like chatGPT?
I grade a lot of papers and encourage/teach chatGPT use. It is so easy for me to detect poor usage. Quality is still easy to distinguish. Skillful use of these tools is a meaningful skill. In fact, it usually requires the same underlying skill! (Close reading, purposefulness, authenticity, etc)
I love chatGPT because it obviates stuffy academic writing. Who needs it. Be clear and direct, that’s valuable!
The goal of a paper, I thought, is thinking and not writing. Without outside help (including the AI), clear and direct thinking is what leads to clear and direct writing. What is being achieved here?
It’s why it’s often not a time saver for writing once you know what you want. But it can help you get going when you don’t. Many other benefits. It does not guarantee good writing, far from it!
> It’s why it’s often not a time saver for writing once you know what you want. But it can help you get going when you don’t. Many other benefits. It does not guarantee good writing, far from it!
Thanks. I was thinking that I need to know clearly what I want, and if I do, ChatGPT would only slow me down. Your perspective makes much more sense.
>All AI-generated text detectors aim for accuracy, but none are perfect and can have multiple failure modes (e.g., Binoculars is more proficient in detecting English language text compared to other languages). This implementation is for academic purposes only and should not be considered as a consumer product. We also strongly caution against using Binoculars (or any detector) without human supervision.
Not very promising, though.
This effect is described in the article. And depending on the context, it can be a feature rather than a bug. If you are using an LLM detector to check if a news article or student essay is "legit", then not only you don't want something from a LLM, but you don't want copy-paste plagiarism either. So for the purpose of checking legitimacy of supposedly original work, then it is a desirable kind of false positive.
I suppose simpler techniques can then be use to check for verbatim copies of famous text.
It’s still not a nuance that most people trying to identify AI will respect even if they know it. Given that constraint, I really doubt the accuracy metrics as well.
We've been looking at the end result and making conclusions about the journey - and that will always comes with degrees of uncertainty. A false positive rate of 0.01% now probably will not be applicable as people adapt and grow alongside ai content.
I wonder if anyone's working on software that documents the journey of the output similar to like git commits, such that we can analyze both the metadata (journey) & output (end result) to determine human authenticity.
If you use edit history, you'll get lots of false positives.
Concerned about this issue I would also run the corresponding outputs through any LLM detection programs I could find (ZeroGPT, etc). None of the outputs have ever been detected as being machine generated.
Wouldn't it need additional data, such as actual proof, to become an allegation or even a claim or charge?
Yes but that hasn't stopped people before.
A false positive rate of 1/10000 would be almost the worst case in fact: truthy enough that people believe that it works, while still creating a vast absolute number of false positives.
What the computer surfaces as an accusation becomes fact by the users of the computer.
Even when the consequences are serious, this happens, an example from yesterday: https://www.nbcnews.com/news/us-news/man-says-ai-facial-reco...
Why should we assume this would go any differently when the consequences are not as serious?
The problem with computer systems flagging things is they are taken as truth.
Another comment on another post perfectly states this human behavior: https://news.ycombinator.com/item?id=39118716
Humans shouldn't use simple heuristics like this to make such serious accusations either way.
At any rate, iMessage does not verify that what is said is true. It just weeds out obviously hostile communication.
The only reason to care is that the implicit proof-of-work signal has broken because LLM text is so cheap. Open forums might need to be pay-per-submission someday...
It feels like we're happy to take the first as a surrogate for the second, or at least being good at the first drops our guard on questioning the second.
We don't need ML-generated-text detectors. We need BS detectors. If they have false positives and trigger on human-generated BS, that's just a great side-benefit.
"This comment could have been written by an AI" is good enough reason to exclude it. That which does not need to be said, need not be said.
As that XKCD's title text says: "And what about all the people who won't be able to join the community because they're terrible at making helpful and constructive co-- ... oh."
That being said, automatic detection seems like a lost cause.
Also we can use more than just the text output. A human writer doesn't generate a piece of text in one pass. Instead they go through the drafting and editing process. We can design devices to capture keypress or pen stroke(iirc ther were studies on fraud detection based on keypress patterns/mouse movment). One can attempt to train a new model to mimic themselves so we need to somehow make sure that the amount of training data required is too much to be worth the effort.
For downstream testing, the goal isn't mainly to verify whether a piece of text is AI genereted but to make sure that a student who can pass the test would essentially have to know the material sufficient well(so this defeats the purpose of cheating).
And it should be illegal to return machine-generate text in response to a discovery request.
Request: "Disclose all documents written between apr 1 and apr 12 regarding topics x,y,z."
Response: "There are 12 billion documents matching those parameters. Here they are."
Request: "Please respond with all documents related to X,Y,Z".
Response: "There are none."
I guess the full-employment economy demands much of us.
What would be an acceptable false positive rate for something like this to be used at schools and universities?
Like, obviously 0.01% is not acceptable, but what would be?
However given that we already have professors literally failing people by just pasting and asking chatgpt, I'm not sure I'm comfortable with that.
It needs to be <= 0.01% false positive for each individual author. If it’s just that 0.01% of all tests are false positive, that leaves the possibility than a given individual might have anywhere up to 100% false positives.
You might find for every 10,000 tests you get around 151 targets. For each one individually you can say they probably did it, but cumulatively you can say that one of them is probably innocent.
Consider using two positives. Do schools and universities generate 100,000,000 essays per year? Sure would suck for that innocent person tagged as having a one in a hundred million chance of not being guilty.
No, our justice system is probably not that effective either. But that is also a problem, so why should we create more problems like that instead of fixing the ones that we already have?
TurnItIn specifically is horrible and should have never been a thing
If the tech froze at its current state, this would be useful for schools. You don't need to expel a student right away after finding a match, but it is a strong indication that something is worth looking into.
(If the goal is to make students write essays, theses, etc. without an LLM writing it for them.)
This is one of the "I don't think that this is the path that we should be taking."
When I was in school, my parents would proof read the essays to catch the spelling and grammatical errors that were in what I wrote (Bank Street Writer had a rudimentary spelling checker but that was it - https://en.wikipedia.org/wiki/Bank_Street_Writer ).
While my parents are both native English speakers and college educated, some of my classmates had less involved parents, or parents that didn't have the same degree of proficiency for writing. Did their essays suffer from a lack of parental proof reading?
In the past few months I wrote two short works of fiction as lore for a game that I play. I used ChatGPT to act as an editor for those works looking at it and occasionally prompting it to help refine a passage.
https://chat.openai.com/share/204de7f7-9cd7-4c45-aa2b-556791... for part of the editor session with it.
Having ChatGPT act as an editor (not text editor but as a critique of the text) helped refine the text that I wrote.
Working with ChatGPT as a tool (that is far beyond the red squiggles in a word processor) to help people working with the written word is a good and useful endeavor. This isn't trying to have ChatGPT supplant human creativity but rather help the person communicate more clearly.
---
I am leaving this in an unedited form, but here is this post with ChatGPT as an editor as an example of how I believe students should try to interact with it. https://chat.openai.com/share/d891f9ac-923b-47a8-8de9-ab7301...
Still, there is a reason why calculators are not used since the very first grade -> it makes sense to learn how to do basic calculations without them.
every time i ask a question to an llm it spits out a generic response format:
'''
well, subject x has a lot of nuance filled with even more nuance. and it may be that x is true but y could also be true, here's a list of related sentences:
1. subject 1 is pretty broad in scope but applies to the question
2. subject 2 is more niche conceptually and applies to the core of the topic without addressing every aspect of it
3. and the list goes on
'''
this is the technology you can't surpass?
I doubt that since ChatGPT trained all other LLM.
> Our approach, Binoculars, is so named as we look at inputs through the lenses of two different language models.
How is LLM generated data out of domain of LLMs? Specifically their github demonstrates with Falcon-7B and Falcon-7B-Instruct models. Instruct models are specifically tuned on their own outputs. We can even say the non-instruct models are also "trained on" LLM outputs as you're using the outputs in the calculation of the cost functions, meaning they see that data and are using that information, which is why
> Unsurprisingly, LLMs tend to generate text that is unsurprising to an LLM.
Because they are trained on cross-entropy which directly related to perplexity. Are detector researchers really trying to use perplexity to detect LM generation? That seems odd since that's dependent on the exact thing LMs are minimizing... It also seems weird because the premise from the paper is that human writing has more "surprise" than that from an LM, but we're instructing LMs to sound more human. Going about detection this way does not sound like it would be a sustainable method (not that LLM detectors are reliable and I think we all know they frequently flag generic or standard text, which of course they do if you're highly dependent on entropy).
=== History ===
First example I'm aware of is the "one-shot" case from[0] (2000) and abstract says
> We suggest that this density over transforms may be shared by many classes, and demonstrate how using this density as “prior knowledge” can be used to develop a classifier based on only a single training example for each class.
Which we can think of as taking a model and fine tuning (often now just called training) with a single epoch, relying on the prior knowledge that the model learned that is general to other tasks (such as training on cifar-10 should be a good starting point for classifying lions).
Then come [1,2] in 2008. Where [1]'s title is "Importance of Semantic Representation: Dataless Classification" and [2] (from Yoshua Bengio's group) is "Zero-data Learning of New Tasks".
[1] trains on Wikipedia and then tests semantic classification on a modified 20 Newsgroup dataset (expanded labels) and Yahoo Answers dataset and is about the generalizability of the embedding mechanism cross domain. Specifically they compared Bag of Words (BoW) to Explicit Semantic Analysis (ESA).
I'll just quote for [2]
> We tested the ability of the models to perform zero-data generalization by testing the discrimination ability between two character classes not found in the training set.
Part of their experiments includes training on numeric character recognition and testing on alphabetical characters. They also do some low-shot experiments.
[0] https://people.cs.umass.edu/~elm/papers/cvpr2000.pdf
[1] https://citeseerx.ist.psu.edu/document?doi=ee0a332b4fc1e82a9...