Additionally, think of cases like paying a lawyer or an expert for an extensive report or opinion on something. Wouldn’t you want to know if that is actually their carefully assembled professional assessment rather than the output of an LLM prompt?
Additionally, think of cases like paying a lawyer or an expert for an extensive report or opinion on something. Wouldn’t you want to know if that is actually their carefully assembled professional assessment rather than the output of an LLM prompt?
How do you think you will know the text was AI generated but the person trying to deceive you won’t be able to undo the watermark?
This article says any AI watermarking can be defeated trivially, by running the text through a tool which strips weird Unicode characters (fine) and then runs it through an LLM which replaces all words with similar words (whaaattt?). Running text through another - probably much lower quality llm which scrambles the words you use would make slop even sloppier. And it’s another whole step you have to know to do. A great many people who use LLMs to avoid doing work will not know about these extra steps, or not make use of these sort of tools.
I would say that's not necessarily the case unless there are zero false positives. In fact, your university situation is exactly where a detection system that works 50-90% of the time would be a nightmare if a meaningful share of the 10-50% errors were false positives.
EDIT: Sorry, I misread your comment above. Yeah, hopefully orders of magnitude fewer kids than the number who are getting falsely accused of cheating with LLMs now.
Increasing the accuracy of these systems - both in terms of false positives and false negatives - seems like a good thing.
If some kids are false positives (detected as using ai but didn’t)then how are they lazy?
A proper watermarking system should be able to have an arbitrarily high accuracy - as many nines as you want. And it should be able to actually report the accuracy of its judgements.
If you're worried about kids being falsely accused of cheating using LLMs, you should be cheering on these developments.
Probability of a false positive is 3 x 10^-5 in one example.
Seems like about 3 collisions per 100 million pictures. If everyone have 1000 pictures that is 3 collisions per 100 000 users.
Raiding 2 people in my city for made up CSAM pictures would be way to high false positive rate.
But you can also make any picture match a CSAM hash by adding picked noise.
Im very confused how this is even supposed to work at face value.
1) If the verification can be done by anyone, then anyone can bypass it.
2) If it can only be done by anthropic then the government or whomever has the special privilege (not everyone otherwise this is just #1) has to make a specific request
Is the point of the legislation to accurately classify text in general or to simply detect the true positives?
Why not both? Stenographically fingerprint llm output to make cheating risky. And use other approaches too.
Most people have no idea open source LLMs exist, let alone how to use them. Right now, they’re much worse than the frontier models.
What about false positives? Imagine being a honest student and then the software declares your work to be AI-generated. How do you defend against that claim? The detection software is a black box and is likely running as a cloud service, so that you have no realistic options for reverse-engineering the false positive detection.
Nothing will save us. You can't automate trust.