ChatGPT Citing Plagiarized Versions of NYT Articles on an Armenian Content Mill
futurism.com
futurism.com
It would work a bit similarly as llama.cpp's grammar functionality, which permits it to produce only sequences matching a grammar, but in reverse.
It of course remains to be argued if this is still some kind of plagiarization. I mean, it would know exactly what to do to avoid being a direct copy..
Machines aren't people, and workarounds based on the affordances given to people should not work for machines, because machines should not be given those affordances.
1) Access to all the internet's data without restriction, including all books and research papers.
2) The model isn't controlled and owned by a for-profit company, it has to be an intra-country initiative run for the benefit of mankind.
Anything else is a wild misuse and bastardisation of AI. We're already seeing all the issues inherent in our current approaches.
I don't see right now how this battle could be won in theory.
There's going to be bunch of content farms, plagiarisers using their open source LLM automation tools, which essentially launder the information from other sources. Even if you denylist them, it's arbitrary for them to spawn further instances.
Eventually the only way would be to not use the Internet at all, and only use allowlisted sources. But that's quite ridiculous to me.