I'm skeptical that anybody generating LLM text is really all that concerned about optimal word choice. Or even particularly good prose. But let's pretend that person exists.
If that person tried, say, an open model and that same model with watermarking applied, I'd be eager to hear their thoughts on the prose quality. Especially if they built an experiment harness and rated a few hundred blinded examples and found a measurable difference.
But getting this upset in advance of any demonstrated problem? It really seems to me like the point isn't the point
The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.
"Which diamonds are shinier, the blood diamond sourced ones or the ethically sourced ones?" ... that's not the same question as "which diamonds are blood diamonds" (to employ an extreme analogy)
Concluding that no one could detect which ones were blood diamonds because they were "equally shiny" is not really correct now, is it?
And the Daring Fireball article does complain that watermarking will reduce quality. If that's what you're trying to check, "which is better?" is the right question.
Also, depending on what I am asking it, I often don't want to use the thumb down or up, as this may mean my conversation is going to have some kind of human review and depending on what I am asking for, I may not want to bring attention to my stuff.
From what I read, two big problems for long-running columnists are getting tire of the work and running out of things to say. I have no knowledge of Gruber, but I can certainly see why people who are expected regularly to have something to say would turn to "AI". It doesn't get tired and is always ready to spew infinite words.
Yes, but that's neither surprising nor a reason to dismiss the anger. People get angry about DRM schemes in video games, even if the slowdown these cause is practically imperceptible. They're angry -- and Gruber acknowledges that factor too -- because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner.
It might be another instance of consequentialism vs. honor ethics. Many consequentialists don't seem to understand that something that doesn't have demonstrable consequences can still have moral implications.
Any potential "slowdown" doesn't even come close to making the list of top reasons people get upset about DRM.
> because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner.
It's LLM output! It's not your domain, it's the LLM owner's!
Yes but thats a thing that degrades something in a catastrophic way, as in I can use the thing one day, and not the next.
A different randomisation system on something that is a text generator which is designed to be unperceptable sounds like the people who are annoyed at FLAC vs MP3[1]
Done right you won't know the difference, done badly and you will.
[1] ex audio engineer, try me.
Good luck explaining that one
What’s the story there? I didn’t know that was a thing and I’m curious to learn more.
Flac takes a raw .wav and effectively zips it up to shave off a certain amount of space. (there are nuances, I think the compression scheme is designed for streaming.)
mp3 is perceptual, so throws away the stuff that humans can't hear. This yields a much smaller file.
However its all a sliding scale like PNG vs jpeg.
a .jpg with a quality setting of 85 will be almost identical to a .png in visual quality. However if you then edit that jpeg, the image degrades and you start to see artifacts. (hence why memes look like shite as they get older)
Its the same with mp3s if you compress the hell out of them, say 64kbit or lower adaptive, then you'll start to hear the tell tail "schlop" noise of mp3-like compression. You might notice it most with cymbals in drum kits. cymbals are wideband noise. as in there are loads of constituent frequencies so if you remove some of the "hidden" frequencies you tend to notice, so they sound more metallic, ironically.
But, all of this is solvable, 256+kbit is more than enough, bonus points for higher sampling frequencies. (however you need a decoder that can actually do that sample rate...)
To the extent that it's anybody's, it's either Anthropic's (they run the service) or everybody's (in that we created the content it's remixing). Legally LLM prose isn't copyrightable for good reason.
that's been his thing since it was just a blog about apple product speculation and update. It's always been tedious.
Taking you at your word, though, I'd be interested to see what you think of the watermarking technology in a blind A/B test.
Do they also do it with code? Do you think deliberately picking tokens that are not the highest probability in code is acceptable for the consumer?
I'm skeptical that such a thing as a universally optimal translation exists in cases beyond the trivial. But if it does, I see no reason to think LLMs are anywhere close to it, so I think nobody will be able to tell the difference with watermarking.
That's certainly true for code. LLM code is at best mediocre. There is oceans of room to subtly watermark generated code without practical impact.
I think in this case it doesn't help that there are multiple watermarking schemes, and the easiest for people to understand is the red/green scheme by Kirchenbauer et al. (https://arxiv.org/pdf/2301.10226), which does technically distort the logits (but I'd argue only in cases where you wouldn't notice it anyway).
I wasn't aware of this gumbel softmax scheme, it seems you're referring to https://simons.berkeley.edu/talks/scott-aaronson-ut-austin-o... ? That's really clever as it doesn't even distort the logits, basically cryptographically indistinguishable from a "real" random sample unless you have the key.
The actual scheme Claude uses seems to be neither of those two though, they say it is SynthId-text which seems to be tournament sampling based.
Prove it, then? It's not a claim that GumbelSoft paper makes: "Regarding generation quality (perplexity), GumbelSoft shows relatively low perplexity"
Cognitive surrender.
The "problem" is that seeing the watermark doesn't mean that the person claiming to be the author didn't make extensive changes to the output of the LLM, or that the LLM wasn't simply the final editor of something that the author had put a lot of work into.
> Cognitive surrender.
I don't know what this means. It's just drama. Don't let the LLM write for you and this is not a worry. I'm not worried about the poetry of LLM output being subtly adulterated.
How does that follow? AI-generated text is already not a perfect emulation of human writing. There's lots of room to affect it laterally without changing the level of quality.
As I understand it, LLMs with temperature >0 can select from many possible outputs. All they're doing is limiting the possible outputs to ones that contain this pattern. I don't see any reason why the quality of that subset should be lower than average. The very best outputs will likely be eliminated, but so will the very worst.
If you get your random numbers from a cryptographic PRNG, then to notice the difference between that and 'real' random numbers even in theory, means you need to break the cryptography. In practice, your gut feeling about how good some text is won't break modern cryptography.
Sometimes when you're trying to write something, it really seems like the exact words matter a lot. Suggestions made to be more direct or use a more common word here or whatever seem to really impact the thought that you're trying to communicate.
Certainly we've all had times when trying to communicate clearly when the specific words seem very important.
This is not entirely accurate. Sure, there's never a token with 100% certainty, but there are often tokens with 99.9% probability, but this technique of course does not change how such a token is sampled.
Indeed.
It's frankly bizarre to see the assumption to the contrary being made by someone who's been passionately blogging by hand for years, who also happens to be responsible for the notoriously vague, humanistic, DWIMmy Markdown standard.
Can you give an example of something smart John Gruber has said or written? Because I can't think of one, but I can think of many dumb ones.