It seems like it is very much a matter of fidelity.
As mentioned in another comment, LLMs (and most popular machine learning algorithms) can be viewed, correctly, as compression algorithms which leverage lossy encoding + interpolation to force a kind of generalization.
Your argument is that a video wouldn't count as pirated if the compression used for the pirated copy was lossy (or at least sufficiently lossy). The closest real world example would be the cases where someone records a the filming of a movie on their phone then uploads it. Such a copy is lossy enough that you can't produce anything really like the original, but my most definitions is still considered copyright.
You would never use a human to backup your financial reports, but the human might be able to give a good overview. You would never use an LLM to backup your financial reports, but they might be able to give a good overview.
AI training data is disposable. There is nothing that could be called a compression algorithm that disposes all of the data you put into it. AI uses training data as examples of what the next token in a token sequence is. The examples are disposable reference points, not the model itself. That's how you get image models that are 20GB in size despite training on 20PB of data. It's 20PB of examples used to form the shape of a 20GB model. You could show it 5GB of training data or 500EB of training data and it would still be 20GB - because it is not a compression algo, it's a 20GB shape formed by external data.
You can compress 20PB of text to 20Gb or even less, if input is super repetitive. So the same with images, if 50% of the images are cats then you learn how to represent the cat pixels with a few vectors and then you could represent all the cats int he world doing all possible cat actions.
But please have the courage to respond to this, when the AI is caught regurgitating the exact text from a popular book, the exact verses from a poem, the exact code function from some code , then how can you defend that is not memorizing things? If a human uses my poem(after they read it) and signs his name under it would you defend them?
And yes LLMs can recall exact material, but it is excerpts and fragments. There is statistical significance to it's ordering. Humans readily do this too (excerpts and fragments), most artists can draw a batman symbol (but not an episode of batman). That doesn't in anyway mean that artists should not be allowed to ever see a batman symbol. It means that artists shouldn't be allowed to get paid to draw one. And they are not. And LLMs are not exempt either.
But the fix is output filtering, just like everything else that can violate copyright. Which is already being done (albeit poorly, but way better than 2 years ago), the same as artists will not draw the batman symbol for you despite being able to.
maybe even simpler, I create a zip format where I randomly replace words with their synonym, or group of words with something equivalent. Would you defend this as original ? why my random transformations are not original while mathematical transformations you will defend ?
And how can you suggest putting output filters to protect only the giants for copyright and everyone else gets screwed.
We should be building robust copyright filters and everyone should be able to contribute their work to it.
but that is a different issue than whether or not an LLM is legally allowed to view a work that is publicly available.
Again, pretty much every artist is capable of off-hand copyright violation on the spot. This has been true forever. We don't bar them from seeing art to prevent this.
We are not talking about the scenario where LLM does a web search and "transforms" your blog as a response to a user prompt.
We are talking about people writing code to torrent content, including copyrighted content as books and newspaper articles then this same people create soem software/script/blackbox and they put this copyrighted stuff inside and this same people also put a filter in the output to attempt and block SOME of the copyrighted material to be spitted out(they did it for Dune, Harry Poter , probably other popular stuff but for sure they did not done it for Romanian copyrighted material or some small writer).
If I publish a poem I do not give implicit permission to anyone to use it as they see fit, and have scripts/software/backboxes train on it and spit parts of my poem out.
I'm sorry, but this a fundamentally incorrect view of machine learning (including, but not limited to transformers).
From an information theoretic perspective the two are essentially identical with the exception that standard compression algorithms do not have a proper "loss" function other than just trying to minimize reconstruction loss with the resulting compression size.
Here's a link to the section on the Wikipedia for more information if you'd like [0]. MacKay's Information Theory, Inference and Learning Algorithms is the standard full text treatment of this topic [1]. Ted Chiang's article "ChatGPT is a Blurry JPEG of the web" is pretty good "pop sci" exploration of this topic if you don't want to get too into the mathematics [2].
0. https://en.wikipedia.org/wiki/Data_compression#Machine_learn...
1. https://www.inference.org.uk/itprnn/book.pdf
2. https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
Humans are totally capable of data compression. This will just devolved into a semantics game of what a data compressor is.
LLMs were not developed to be, do not function as, and are not use as data compression utilities. Please, come knocking when a service provider exists that will use LLM's to compactly store your company data.
https://arxiv.org/pdf/2309.10668
Transformers are also used in the top algorithm right now on the Large Text Compression Benchmark. https://bellard.org/nncp/nncp.pdf
Again, from a information theoretic view point, this is exactly what they are doing, how they where developed and how they function.
I don't know any serious researcher in ML that would find this claim even remotely controversial. It's really not just "a semantics game", its a part of a foundational understanding of the topic. If you want to understand LLMs from this perspective, a good place to start is with an auto-encoder which does try to learn a standard compression algorithm, the move on to more sophisticated embedding models (found in a lot of recommender systems) which try to learn an additional objective on top of minimizing reconstruction error. You'll then see that Transformers and all other major NN architectures fall out of these basic principles.
> Please, come knocking when a service provider exists that will use LLM's to compactly store your company data.
This is literally what every vectordb company does right now, as well as all "chat with your docs" type startups.
Can you get transformers to regurgitate information verbatim? Yes.
Would anyone in their right mind rely on a transformer to do so? No.
Would anyone in their right mind rely on a vectorDB to do so? No.
Would anyone in their right mind use a vectorDB/RAG/SQL/transformer combo to do so? Yes.
Is youtube going to drop VP9 for GeminiEncode to save google billions in bandwidth? No.