But of course, watermarking or checksums stop working once the general public runs LLMs on personal computers. And it's only a matter of time before that happens.
So in the long run, we have three options:
1. take away control from the users over their personal computers with 'AI DRM' (I strongly oppose this option), or
2. legislate: legally require a disclosure for each text on how it was created, or
3. stop assuming that texts are written by humans, and accept that often we will not know how it was created
[0]: Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., & Goldstein, T. (2023). A watermark for large language models. arXiv preprint arXiv:2301.10226. Online: https://arxiv.org/pdf/2301.10226.pdf
Also, technical enthousiasts will run LLM's locally, like with image generation models.
In the long term, when smartphones are faster and open source LLM's are better (including more efficient), I can imagine LLM's running locally on smartphones.
'self-hosting', which I would define as hosting by individuals for own use or others based on social structures (friends/family/communities), like the hosting of internet forums, is quite small and it seems to shrink. So it seems unlikely that that form of hosting will become relevant for LLMs.
Fist, if it should work, you'd need fuzzy fingerprints. Just changing a linebreak would alter the SHA sum.
Secondly, why?
I generate some text using ChatGPT.
ChatGPT sends HaveIBeenGenerated a checksum.
I publish a press release using the text verbatim.
Someone pastes my press release into HaveIBeenGenerated.