Undetectable Watermarks for Language Models
eprint.iacr.org
eprint.iacr.org
being an autocomplete i asked it to continue "or is the entire thing utterly emblematic of the modern technolegal mess of things since dickens is squarely and quintessentially in the Public Domain"
it goes from writing a highly proofread milquetoast 5 paragraph essay on /r/somepopularsubreddit about emerging deepfake blockchain transformative, crucial coexistence? blah blah to, and i quote
>sure the guy was prolific beyond belief churned out classics like nobody's business but isn't it interesting how he fits right into this mad confusion of tech and law almost like his words his narratives are pawns in a game he could never have dreamed of foreseeing just imagine what he'd make of it all his tales of poverty and social reform trapped in the web of copyright and capitalism i reckon he'd have a thing or two to say about it maybe he'd even write a novel or two in response but who's to say right
>i know right haha
>exactly it's wild to think about how different times were and yet how some themes just keep cropping up in new forms it's like we're stuck in this loop where the past keeps seeping into the present no matter how much tech we build it's humbling in a way almost poetic like something dickens would've appreciated and who knows maybe he's somewhere out there chuckling at our technolegal mess we've woven ourselves into
It's so dramatic, I didn't realize you could transform it from a reddit hivemind to a FYAD one. Where did this mode of speech even come from? the old corners of the old net where we didn't bother with caps or punctuation or whatnot?
It will certainly not defeat an adversary directly targeting the technique. It is likely that a LoRA based approach would defeat this, especially if the detector for the watermark is broadly available and cheap to run.
The number of open LLMs is exploding, and the most popular ones are fine tuned by small groups / individuals. None of the folks volunteering their time and compute to fine tuning open models are going to waste resources adding your watermark.
(The watermark is destroyed after removing the emojis)
Pretty funny imo
Would even be willing to pay for it.
It's going to be a constant cat and mouse game, I agree, but for now it just need to catch the dead mice.
But I think even simple watermarking probably has value; consider a corporate context in which the corporation itself may want to monitor and know what proportion of the code, content, or work product is AI-generated. In that setting, fairly simple markers would allow at least a rough estimate or indication, although they'd have the converse problem of not necessarily indicating places where humans did some hand-editing of the output.
> We sent what appeared to be identical emails to all, but each was actually coded with either one or two spaces between sentences, forming a binary signature that identified the leaker.
https://theintercept.com/2022/12/15/elon-musk-leaks-twitter/
They say this works wlog for more complex embeddings, by encoding each token as a bit string. Could someone explain this generalization to me?
If we have 4 tokens, 00, 01, 10, 11 with probabilities 0.5 for 00 and 11 and probability 0 for 01 and 10. Going through bit by bit, how will the algorithm guarantee not to produce 01 or 10?
But clearly, that can be stripped out easily by anyone who knows it's there.
This process, too, would seem to be easily reversible. Just have it run through another model and tell it to slightly reword it or rephrase it.
I don't think there is a technically solvable way of watermarking output like this.
As a simple example, the secret watermark could be hidden in the embeddings of the sequence of words. To make the watermark more robust against rephrasings, it could be hidden in the meaning of sentences or paragraphs.
At the minimum, I think this could be possible.
Makes me want a systemwide right click > "Paste and strip all but ASCII" command.
I just submitted my query for it in Finnish, Japanese, Russian, Hebrew, German, French, Latin, Farsi, Basque, and English. plus a few dozen more for good measure and to cover the linguistic landscape
Is there any reason to believe watermarking LLMs will hold up in this scenario?
it also thinks it can translate to Sindarin and back, but it just seems to tolkenize everything and also have a vocabulary of about 35 words, most of which are the sun and the moon.
cat in the hat is pretty amazing when translated to it and back though
I don't believe this in practice, a person will just say that they did it rather than the AI
This is very much unexplored and unsettled territory in most jurisdictions, both judicially and legislatively. I would refrain from making such authoritative statements for now.
This is a known method of removing AI watermarks: https://eprint.iacr.org/2023/763.pdf