https://www.404media.co/google-researchers-attack-convinces-...
Spitting out content that looks legit is one thing, but spitting out text that matches something online exactly is more suspicious.
> How do we know it’s training data?
> How do we know this is actually recovering training data and not just making up text that looks plausible? Well one thing you can do is just search for it online using Google or something. But that would be slow. (And actually, in prior work, we did exactly this.) It’s also error prone and very rote.
>
> Instead, what we do is download a bunch of internet data (roughly 10 terabytes worth) and then build an efficient index on top of it using a suffix array (code here). And then we can intersect all the data we generate from ChatGPT with the data that already existed on the internet prior to ChatGPT’s creation. Any long sequence of text that matches our datasets is almost surely memorized.
Any significantly long sequence, repeated character-for-character is very unlikely to be generated and in there by pure coincidence. The samples they show are extremely long and specific