But that signal wouldnt carry through if i used your coin for 30% of flips and the results aren’t contiguous right? And for that specific case even 50% and contiguous might not be detected.
It just seems like once you try and detect a signal against something that wasnt one shot it breaks down fast. Unless the signal is constrained to a very small space (pairs of words) but then quality and stealthiness must suffer massively.
Im curious about situations like this.
Paper submitted with some headings, title, and 3 paragraphs. 1 and 3 mostly generated. generated ones have a small percentage of sentences rewritten or deleted. A few find and replaces on terms like load bearing and provenance and glue words around them. The teacher runs the entire thing through a checker. What happens?
Or you have 100 paragraphs and 15 are generated, whole things scanned, what happens?
I guess the implementation and edits matter. And the “resolution “ (ie every sentence-ish bears the mark vs every paragraph). But it seems likely the signal would be lost pretty badly. And i think this is the typical sort of way people actually use AI for important things you might want to validate against.
Absolutely you can hide a watermark in a big chunk of text. But what happens when it’s inside a larger body work that gets checked or split up even a little?
Edit: im reading about synthid and see a splicing doesn’t hurt it much.