The Science of Detecting LLM-Generated Text (2024)
dl.acm.org
dl.acm.org
Which is really boiling down to text having statistically very similar properties to human generated one. Introduce a more motivated attacker and the text would be indistinguishable from real (with occasional typos, no use of "delve", "it's not x its y", emdashes and so on).
It really is a lost battle: you cannot embed extra information in the text that will survive even basic postprocessing (in contrast to, say, steganography)
You've just described a “base models” (or pre-trained model), but later training stages (RLHF, GRPO, whatever secret sauce model makers use) induce a strong bias in the output.
Also, being “statistically identical to human generated text” doesn't mean it's unrecognizable, because human generated text exhibit many various clusters (you're not texting your friends with the same language you're writing a book with) and an LLM can, and in practice, do, use language that is not appropriate for the tone a human expects in a certain context (like when bots write LinkedIn-worthy posts in reddit comment section). The “average human-looking text” is as unnatural to us as a “synthetic average human” with one testicle and half a vagina would be.
The top models are also the latest:
Gemini 3.1 Pro: still a bit of a gremlin, but will probably stay on top until the other model makers go xkcd 810 and target this benchmark
Gemini 3 Flash: current favorite of writers using it as a helper for its speed and decent prompt following
https://books.google.com/ngrams/graph?content=delve&year_sta...
Mirroring real human text is only the basis of training. Afterwards they get aligned a.k.a. lobotomized.
Once you give the llm examples of your prior work and ask it to continue its style its game over for detection.