I think that was intended, yes.
I think that was intended, yes.
It's still possible to use Claude to proofread - highlight grammatical, flow, structure, logic errors and make simple suggestions for you to pick and choose or adapt as you wish. No watermarking will flag your text. No flaw accusations of LLM authorship will haunt you. All will be fine.
But if you want an LLM to rewrite your text, that's (a) not proofreading, and (b) should be flagged as LLM generated ... because it is.
"When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to. Depending on the length of the text and how heavily Claude has edited it, those changes might not be enough to make Claude’s involvement detectable."
(emphasis added)
A proofreader, like an english teacher, can return you your writing simply annotated and marked up, with suggestions and edits in red pen for example.
Then you rewrite your draft into a final using those edits and notes as suggestions.
Making an LLM act like that proof reader would likely not cause your output text to be labeled as generated, even by this system.
You could use your own hypothetical house elf to do it for you, or pay someone to do it. LLMs are just cheaper for a certain set of problems.
People will find ways to circumvent this, so this limitation will only hit the technically less adept people.
In the fact that you didn't write it.
> You could use your own hypothetical house elf to do it for you, or pay someone to do it.
Yes, and those would be similarly problematic (and more expensive).
This is a fact and this is generally not a problem. Customer support guy Joe did not write that email to you with a refund: someone else did it and Joe did pick the template. Alice did not write that post card to Bob, someone else did and she just googled some nice text. We deal with a lot of content that wasn’t written by the person who signed it. That content, when written by LLM, may indeed contain watermarks and nobody will care about the choice of words, because only the meaning matters in such communications.
People pay too much attention to authenticity here, which is no more than a demonstration of an effort. LLM can and should write scientific articles because the real effort is in directing research, not summarizing it. LLMs can and should write news, because it is cheap and efficient, and real reporting is in discovery. LLM can and should write fiction and make movies, because there is no reason why creators of various junk should earn their money easily. LLMs do not replace real talent. They just emphasize for an average person how easily replaceable they are. And that‘s ok. Creative industry is a blue collar job now.
In case of the former, there is a problem worth solving.
Factory farming also makes meat cheaper than organic practices. I'd just like to know which one I'm getting.
And if you copy-paste the answers from LLM, I think it's only fair the end result gets flagged. You're not writing it yourself.
I wonder if it would even get flagged in that case, because wouldn't the probability distribution of a token when the LLM is suggesting an edit to your writing be different than the distribution of that token once it is in the context of the text it's editing?
A point of confusion for me, however: is every watermark unique? Is every algorithm for watermarking going to vary amongst models and amongst model versions? Will each model publisher keep this watermarking as a trade secret, that they alone can detect? If so, this can't scale! How do you detect "JoeBob 4.3 LLM" output? By querying every single model's watermark-detector? And if they all work by re-running the model and using tokens anew? That is extraordinarily wasteful.
If a watermark is not self-evident, or universally detectable, then it is no good. Take, for example, US currency. The security measures are published and well known. Any count-out room in retail has a big poster indicating how you can detect authentic US bills. Nobody has to accept non-US currency in the US, and so the only authenticity you need to worry about is your US bills alone. LLM watermarking has none of this in common. Currently sounding like a shitshow, if you ask me.
Wouldn’t having that be enough to eventually reverse engineer the key?
Removal may come down to changing every third token to a different one.
Working backwards: if it is possible to confirm 100% confidence that a chunk of text is LLM output, then it is "PD until proven otherwise". How can a human reliably assert human authorship of their source text? When all watermark tests fail? Is that proof of humanity now?
If a human proves human authorship, and LLM watermarking tests positive, then is that going to be considered a "derivative work" or not? What if there is an applicable license for the source work, such as "CC-BY-ND" that prohibits derivative works?
This has not been court-tested, and I expect that it will need testing at that level before we can have any assurances.
As for it being "uncopyrightable" if it were the output of an LLM, I think computer people are making a very aspie interpretation of a single decision. I think it's more that the LLM (and thereby its owners) cannot itself hold a copyright on its output, that output has to be touched by a person before it is copyrightable. A particular view from the top of a mountain can't be copyrighted, for example, but a photograph of that view can be.
I'm not sure it at all precludes a "robosigning"* sort of situation, where machines generate output, hired temps sign and claim that output, and immediately sign it over to the people who hired them (as a work-for-hire.) Copyright is stupid, artificial law, not logical.
-----
* https://www.mortgageauditsonline.com/what-are-robo-signers/
For the watermark to be detectable, the text needs to be like 75% AI generated.
If you have an LLM “touch” one section of the article, it’s not gonna be detectable.
Here is some text that is copyright to me. As you infringed my copyright, please pay my $5000 license fee for every user who has read it:
> Copyright law was updated in a very helpful way in the last twenty years sometime so that as soon as you post something to the internet you have copyright. If you need a citation don't hesitate to ask someone else.
They will need to dodge around the EU requirements but it will probably just come down to an alternative method to watermark or a contractual assurance you won't mis-represent the source of the text.
I hope the other providers will add a geographical limitation on this EU rule.
(On a side note, I wish they would replace those EU beauracts with LLms).
but proof reading is a linter, not a writer. the proof reader will say "I think this is clumsy can you try x,y & z"