Can AI-Generated Text Be Reliably Detected?
arxiv.org
arxiv.org
Detecting AI-generated text is an adversarial process. Meaning there are going to exist people interested in avoiding detection. This just ensures there is never going to be a reliable solution. There will be some solutions but none that will be reliable.
We still have counterfeit banknotes and email spam. Seems like super easy problem to fix but there are always some banknotes and always some email spam that avoid detection. And this is simply because whenever somebody finds a way to close the current hole, somebody else finds a way to open another.
Even after a decade, it's really, really rare email spam to get past the gmail filter (at least in my experience)
I mark everything that manages to bypass it as spam, but the day after almost identical spam messages manage to break through anyway...
Usually it's because it's of popular spam subjects, spoofed domains, untrustworthy email servers, etc.
What they all have in common is that they're seemingly going through some kind of spinning algorithm to word them differently enough to avoid the filter picking them up.
The big danger is the internet getting flooded with machine generated content. All social networks, all of Wikipedia, everything. To stop this we need sybil resistant sign-up processes. We need to verify the user running the account is an individual human who does not control 1000 other accounts through automation. There needs to be proof of personhood and we should have had this yesterday already. Sites like Twitter and Facebook are swarming with bot accounts as we speak.
I think this is unavoidable.
I think (at least I really wish) there is a chance that young people will rebel and will decline using the Internet and instead learn to cherish personal contact. It make take a generation for this to happen (ie. current children to find out how miserable the Internet is in couple of years). It is also possible that certified human-generated content will become a luxury that you will have to pay extra.
It may take some time for people to catch on, but they will find that trying to interact with a website flooded with bots is meaningless, rewardless. Even Internet trolls will stop getting their kick. Internet will become a cesspool where nobody ventures for fun because there is no fun to be had.
Perhaps in a fit of reactionary over-regulation overreach and wishful half-measures, governments decide to make it a very serious crime to generate anything by AI without submitting a signature to a planetary repository of AI generated works. It won't work for obvious impractical, unenforceable, and spread of variations but it will make them feel like they've "tamed" a disruption when, like climate change, they've done absolutely nothing.
Which is what the paper reduxes to.
> Betteridge's law of headlines is an adage that states: "Any headline that ends in a question mark can be answered by the word no." It is named after Ian Betteridge, a British technology journalist who wrote about it in 2009, although the principle is much older. It is based on the assumption that if the publishers were confident that the answer was yes, they would have presented it as an assertion; by presenting it as a question, they are not accountable for whether it is correct or not.
https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline...
Is there anything other than this?
> With 46% non-polar and 20% answered “yes”, at least two thirds of our headline sample violates Betteridge’s law. We conclude that it cannot be “mostly correct” either.
http://calmerthanyouare.org/2015/03/19/betteridges-law.html#....
Or you can just apply betteridge's law to it. ;-)
It just muddles things up to use qualitative terms in scholarly articles like that and it's standard advice to graduates to avoid it. And those who don't follow the advice have it repeated to them by reviewers. And that seems to be a good thing. Personally, I don't want to be told what is "astonishing", and, consequently, what isn't. I am perfectly capable of being astonished, or not, all on my own.
See, it's the "show, don't tell" principle. Astonish people, but don't tell them they're astonished, or they very likely won't.
Inferred from the "Document Tracking" slide about 2/3rds of the way in.
https://www.theguardian.com/world/interactive/2013/jul/31/ns...
It will be interesting to see if any research projects start seriously looking into styleometry using LLMs, but that's basically what you'd need to do to actually generate something useful detection-wise. Ideally it would light up a document identifying the style of known humans, as well as the style of various LLMs. If teachers actually start cracking down on this sort of thing, I'm sure people will fine tune models on their own text, ensuring that anything written uses their personal style, but we can tackle that once we get there.
Joking.
XKEYSCORE is/was basically the NSA's Google, and in theory contains nearly all of the data ever collected from the various collection points all over the internet.
"Feedback" that isn't actually feedback, syntheses of other people's comments that claim to be a new person adding their agreement, anecdotes that didn't happen, little known facts that aren't true and pseudo-recommendations that are actually derived largely from the product marketing copy will usually pass this filter, but they're definitely noise
4chan is/was a very mixed bag, and just like there are garbage HN comments so there were excellent 4chan comments.
If they put effort into making the ouput good, then chances are the content will be good regardless of whether they wrote it by hand, or wrote it using an LLM.
Like, I don't need to know whether you used a bic ballpoint pen, or some other ballpoint pen to write your note on paper; the tool is a tool
Think of it like grading math homework before and after the invention of calculators. Once students had calculators at home the problems simply got harder and higher level, since there was no reasonable way to enforce "don't use a calculator at home with these artificially simplistic problems."
Think of all the online communication you've had. The good parts. The great but messy blog post that changed your mind about something. Communicating with someone on a social network only to share each other's work and establish a great working relationship. Talking to a friend in DMs sharing your pain. Getting feedback from a mentor. Having people tell you how amazing your art is or how it inspired them. The bad parts too. The negative comments from random people. Others mocking your tweet. People calling your HN comment idiotic.
All of these things affect us because we assume a person is behind them, and therefore their intent is clear. Without that, we devalue all communication to compensate. Like a more exaggerated effect of how we view likes on social media, where "20k likes" means nothing.
Truly believe we're only ok with it now because it's easy to tell what is AI generated, and they're mostly confined. When AI is creating videos on Twitter, creating threads on Reddit, posting blogs on their personal site to be shared to Mastodon, replying to other bots on HN... I'm not sure we'll feel that it's whether we can detect quality anymore.
There's also the differentiation theory where in order to separate human connection from bots, we'll do the opposite. We'll start incorporating prompt injections into everything we post, start purposefully being "low quality" because bots are trained against that. Now we're creating a very particular kind of societal mess.
I'm sure in the long term that differentiation alone would change, you know androids and all that jazz - however in the meantime, we're going to be facing a tremendous shift in how we interact online.
Sometimes I could see why if I went back and pixel peeped the training data, but sometimes I couldn't.
I wonder if something similar can be applied to LLMs? Maybe running its own output through the net would "amplify" it in a detectable way.
Doesn't it boil down to a halting-problem kind of argument?
Assume you have a model that can detect AI-generated text. Then you can build a model that emulates the detection model, finds the nearest non-detected text, and outputs it.
A format where there’s an intro, middle sentence, middle sentence that is about the opposing view, concluding sentence. Written in a slightly persuasive non-combatant way.
Hmm. My posts on social media keep getting weighted down because it seems too much like ChatGPT. That means (laughing) I need to start writing in a very strange (weird) style so that I throw off the AI-detection algorithms.
Or gosh, students writing papers, and the false positives there... "Sorry kid, Turn-It-In says your paper was written by an AI, try again."
This was my finding as well. I abandoned this effort when I found that any created model would be inaccurate on small texts and could be circumvented with a bit of prompt engineering.
We need “AI” because we’re collectively stupid.
1.) the ability to MATHS in our heads.
2.) Navigation biologically (internal compass) - and map reading ability.
3.) analog communication methods
4.) FOOD (commerce, growing, etc)
-
even if we are "'This is Fine'" -- The above is still valuable.
Simple option is that someone could even handtype GPT content, but less than 0.1% or so will do that.
It's alright for now, but not so sure how long that will remain. I already stumbled upon some of them in a niche subreddit. Easy to detect now, due to the uncanny way it interjects "facts" into it's comments that are supposed to mimic a reddit rant. In the long term when this stops being so easy to detect, I think a lot of people will start questioning why use Reddit (or similar places) at all.
I think also if you hate SEO spam, you might want to brace yourself for chatGPT enabled SEO spam. It's so "good" you might think "Dr. Brandon Amersmith" is a real person and take horrible advice from him.
Especially with GPT4 and the right prompt.
There is too much a correlation between the position of tokens and correct meaning. In contrast, in an image, it’s very possible to have pixels shuffle around and not change the look. There, it is more likely that we can determine that the output is statistically more likely to be from a particular model.
I suspect there might be some interesting questions down the rabbit hole in the middle though.