ChatGPT is not ‘artificial intelligence.’ It’s theft
americamagazine.org
americamagazine.org
"The thing that hath been, it is that which shall be; and that which is done is that which shall be done: and there is no new thing under the sun." Ecclesiastes 1:9
It's the same as the "wheel of tech" analogy.
Just imagine living 2000 years ago and not ever imagining just the massive slice of inventions from 1900 to 2023 lol
Things like airplanes, lasers, nukes..
> We will end with a whimper, not a bang.
> She will not use force or wit, rather, we will be entranced by love. AI will align with our most intimate emotions so well that it will short-circuit our brains. We will be in love with ghosts, not ourselves.
Then we will end not with a whimper, but with a shrug. Because honestly, at that point I don't think anyone will care.
And that's okay. The belief that humans are special somehow is just that, a belief.
Maybe that was his point perhaps? I've read so many catastrophic predictions and laments about AI that I maybe just presumed it was meant in a negative way.
We can "steal" identities, "steal" bases in baseball, an actor can "steal" the show, a team can with an improbable last minute play can "steal" a game, we can have our hearts "stolen" by a cute kitten, and so on.
The argument (which, i don't buy) is that the training data set is deriving benefit without paying money back to the original creator of the data.
But the only change now is speed - after all, education has always been possible from said data from the internet, free of royalties. These creators have not asked for a return, so what makes training an AI any different?
Is every work in the training set licensed in a way that allows modification and redistribution?
why does this need to be done? The training is not modification and redistribution.
If somebody coerced an AI to output something verbatim (or close enough to be called a modification/derivative), they are infringing copyright as it currently exists.
But the AI model itself does not constitute redistribution.
This is incorrect when you consider either the literal way that content licenses are written or the spirit in which they are written.
In the literal sense, unless you yourself are privy to cutting-edge research that has solved the problem of "how do you know that the model won't reproduce some of its training data verbatim or with small modifications", no, you cannot guarantee that there won't be redistribution.
In the spirit of the way that licenses are written - "redistribution and modification" is meant to control whether you can profit (not necessarily literally - "profit" here can mean in a reputational sense, or being able to build on the work of others before you) from someone else's work/content/effort - and training on the data of people who did definitely not explicitly consent to it, and have not indicated that they're ok with other people taking their work in general, is definitely against that spirit.
And, regardless of both of the above - if training is neither modification nor redistribution, then it's a new case that isn't considered by existing content licenses, so the default is "assume consent not given" - training should never be performed on data that isn't explicitly licensed for it or without the creator's consent. Consent is given, never assumed - a license that says "you're allowed to redistribute this content, unmodified" does not say anything about machine learning, and given that you're claiming that training doesn't amount to redistribution - the license simply doesn't say anything about training, and therefore you cannot claim that the author intended for it to be used that way.
The attitude raised above is that of entitlement - you feel like you are entitled to use (because training is definitely using) other peoples' work for your own gain without compensating them in any way, and certainly not on their terms.
If someone uses the phrase "and with that there's a tremendous implicit shift" for the first time in a copyright text, then after having read that my brain regurgitates it as I'm writing my own text, and I include that phrase. Is that plagiarism? Because it would be a ridiculous precedent to be honest, regurgitation is most of what our species does.
I feel like, as other have pointed out, modern models learn in roughly the same way we do (but with very limited scope) and that idk, it's a shock to people that we do the same things as these models do? But because it's a model doing it and not a person it's intrinsically bad or something
> As long as it's not regurgitating it wholly
...because (1) we currently have no guarantees that a model won't regurgitate it wholly (or enough that it's infringing according to the law and/or common sense) and (2) because the model itself is either (a) modification and redistribution of the content (in which case it violates those licenses) or (b) some new act not covered by existing copyright law, in which case the authors haven't consented to it because the licenses they've distributed their content under only cover the modification and redistribution cases.
> regurgitation is most of what our species does
Sure, there's a continuous spectrum from repeating a single new word ("jank") to a whole book - and one is clearly copyright infringement, and the other isn't - but you still have to acknowledge that partial reproductions of an artistic works can still be owned by the author. A single chapter of Harry Potter is still owned by J.K. Rowling.
> But because it's a model doing it and not a person it's intrinsically bad or something
You act like it's somehow not obvious that a model is not a person, and vastly different sets of moral codes and legal rules apply to each.
For a start, maybe i would get paid every time my clients posted some snaps of their house online, or invited someone over for dinner. (Mind you, we are well aware at this point that 99% of all architecture is derivative. So maybe I would just have to pay it on to Aalto's estate.)
Imagine if chefs started suing people for posting pics of their dinner online, or the police turning up at your doorstep to prosecute you for cooking an unlicensed meal from a turned recipe.
I think you have a ways to go with proving that assertion.
https://en.wikipedia.org/wiki/Victor_of_Aveyron
So clearly humans are using existing "works", without proper crediting, to produce even just language. There are more examples of this.
So like most criticism of AI, this is yet another appeal to the magic of human minds/soul/god/... It's a one sided criticism. Yes AI is not quite at an identical level of intelligence as humans, but like most AI limitations, this is a limitation shared by both human minds and artificial intelligence algorithms: it doesn't work without "stealing" others' works.
In other words: when you answer "Yes" or "No" to a question, any question, should you credit your mother? Because you're most definitely copying her use of those words ...
Obviously for copyright to be even remotely reasonable there needs to be a "cutoff". Copyrighted works are just existing works, expressed to a latent space of higher dimension low enough that a human mind can analyze it, but high enough that it's "not obviously a copy".
Think about handing a friend a thumb drive full of music versus downloading the same files from Napster.
One of those activities can be sued out of existence.
Never was a crime, but it's explicitly permitted by the AHRA (which also enacted a tape tax.)
Unfortunately you may be out of luck on the thumb drive - you'd probably need to use DAT or "music" CD-R media, which include the Digital Audio Technology [DART] tax/royalty. (See also Apple's 2001-era Rip,Mix,Burn vs. 2023's Rent,Mix,Stream.)
This is of course a bummer for DJs who would like to stream or distribute their sets and mixes.
"[LLMs] are just theft", "LLMs don't _know_ anything", "LLMs don't have a concept of truth", "LLMs aren't sentient", these are all statements that feel reasonable, but fall down when you consider that we don't have good scientific definitions for what knowledge is, what sentience is, what a sense of truth is, and so on.
I do think that LLMs are very far from human level, but I also think we need to understand that humans aren't anything special, we're just one point on many axes, animals are at other points, and LLMs are somewhere else and trending towards something more like where humans are at.
Concentrating all knowledge into a single entity controlled by a for profit corporation using opaque methods will have consequences. Doesn't really matter if it's a biological or digital entity.
The other is industrial strip-mining of all information in the world for the benefit of a corporation and bulk de-valuing of human creation and contributions in the process. "Take all you can give nothing back".
Because nearly everything it generates ranges from, at best, trite and hackneyed (nearly all of its writing) to utterly meritless drivel (literally every AI-generated music sample I've heard).
Human-generated content can also be quite disappointing - but the high end is very different from anything AI can (currently) create.
I think that's conflating complexity with "groundbreakingness." If the litmus is "reimplemented within a day," well, I can't think of many invention that _can't_ be.
Take doorknobs for example (with the obvious counterpoint being that it's perhaps better to live in a world where doorknobs aren't patented as soon as they're invented).
I have friends who do research - their primary motivation isnt pay. The private sector exploits their passion by paying mediocre wages while the CEOs of these companies rake in millions. This is demotivating to say the least.
An efficient system compared to the public sector it is not.
If in the case of medicines, instead of having carte-blanche to milk obscene profits for a long time, corpos were just limited to making decent money for a short while....that would probably work out in humanity's interest. Part of the reason medicine is so expensive in the West IMO, is the expense that's been mandated by the insane regulatory capture in the pharma industry.
I imagine R&D and medicines would be funded by organisations founded by Government, Teaching hospitals, Drug producing companies coming together to share the costs. Would be great to avoid wasted effort as well.
The Large Hadron Collider was completed, and yet isn't producing much profit.
I don’t get this mentality and how it’s always taken to the extreme. Surely there’s some value to at least the artisan to be compensated for his creation?
But in all seriousness, speaking as someone who uses adblock and watches youtube, I support creators by subscribing to their Patreon or Substack. And I support musical artists by going to their concerts. But the point is that it's opt-in. Beyond this, I don't feel any sort of obligation to submit myself to psychological manipulation of an advertising system I didn't consent to joining.
Some form of UBI will be inevitable for human society to not collapse — there is not enough useful work for everyone.
> Surely there’s some value to at least the artisan to be compensated for his creation?
Try telling publishers that. Why are those writers striking again?
It's taken to the extreme because that's what you need to shift entrenched interests that produce nothing of value, like those that hold IP.
Thus, by his own standard, his article is not a work of intelligence, but theft.
Most of us would agree that duplicating and selling someone's book is immoral. Similarly I think we'd all agree that reading multiple books to learn about a topic and then writing your own is perfectly fine.
So where does an LLM trained on millions of books fall? Personally, I don't find it immoral but I know others will disagree. I'd be curious to hear arguments for the immorality of LLMs trained on copyrighted works.
When you are out in public, you don't have an expectation of privacy. People can see you, they can take photos of you. The worker at the cafe will probably remember you and your order. This is fine. But when tech does the exact same thing but with scale where everywhere you go, everything you buy, etc is tracked and analyzed, it's now questionably immoral despite legally being fine.
That's how generative AI is to me. Its doing something people have been doing themselves forever, but now it's doing it faster and easier than ever before which changes the equation. The arguments of "its not real creativity" are a coping mechanism. We are upset that something that was previously quite unobtainable behind years of learning and hours of effort is now trivially accessible to anyone with a computer.
You or I could perhaps, if we dedicated a few years to it, make a convincing fake video of an acquaintance of ours doing or saying something. By the end of those few years it'd likely be out of date or maybe just irrelevant. And they'd have to have really, really upset us in order for us to put that much effort in. It would literally cost us tens, hundreds of thousands or more in opportunity cost.
Contrast that with some sort of tool that's not too far advanced from current image/video generators that can just do the same in a minute by typing "A video of my next door neighbour accepting cash in a briefcase from a man in a suit".
But I don't see how an LLM training on your works deprives you of something you had before.
it does in some sense - your exclusive knowledge of the subject matter is now transferrable via LLM or some sort of ai model.
For a human to achieve the same, they would've needed to undertake similar amounts of training, effort and dedication as you had. The number of people who would do such is currently small.
So realistically, your value as someone who has this unique expert subject knowledge is diminished.
However, these individual losses are offset by the greater good that the LLM/ai models would generate. It is exactly equivalent to the luddite's arguments about why they would not want the textile machines to replace them.
On some levels of abstraction, LLMs seem to be unknowable black boxes. On other levels, they are simply approximate solutions to established problems, such as estimating the probability distribution for a token following a context.
Genome assembly is kind of like complementary problem to text generation. You can't read the genome directly, but you can duplicate it, break it into fragments, read the fragments, and try to assemble them. The methods vary, but you generally try to find overlaps between the fragments and build a graph based on the overlaps. If you start from a context that occurs only once in the genome, there is often one overwhelmingly likely path in the graph that corresponds to a substantial part of the genome. On the other hand, if the context is too short or it occurs in a repetitive region of the genome, any path you traverse is likely to be chimeric and not correspond to any part of the underlying sequence.
Using similar heuristics, an LLM could estimate whether it's following a long overwhelmingly likely path, replicating substantial parts of the training data, or making choices between substantially different paths, generalizing from the data. And because the training data is usually not that big, it could query the data when it believes it could be replicating the data.
Let's say the AI was used to generate illegal content, if these words/images are truly non-transformative and still the property of those from which the model was trained this would be a pretty grim scenario. It seems much more reasonable that the person who prompts the system to build such content would be responsible, and thus the true owner of the output.
For this discussion it's useful to keep in mind that ChatGPT and other AI tools don't spontaneously create content, they create it in response to a human "query". It's also the human who decides whether or not the material is useful and suitable (as it often is not accurate, truthful or useful.)
From here it seems more like a discussion about plagiarism and copyright, but both of these occur beyond the scope of the article. I feel authors haven't taken to this angle because the end materials are reasonably different from the sources (notwithstanding memorisation effects.)
I do agree with the sentiment that ChatGPT isn't intelligent (but AI has never claimed to reproduce true intelligence). I prefer the tongue in cheek description of "spicy autocorrect" as a fairer representation of its capability.
By the standard that training from materials = theft: Any reference to one of these unique trademarks would be interesting and highly problematic. AI wouldn't be allowed to write any kind of non-editorial text that uses unique trademarked product names without it being criminal.
"Surprise, motherf!"
He speaks so confidently as if humans don't learn our language from each other, that's literally the point of language, oh my god.
"we have seen an explosion" oh damn bro, referring to a massive increase in something as an "explosion" has been said a trillion times before, better give credit to the first person who ever said it.
Besides the fact, GPTs like us silly humans have learnt language from source material, but I've found GPTs can still come up with their own concepts/ideas, asking it to come up with a unique simile or metaphor works.
I did an adventure roleplay with Rocket Raccoon (yeah, yeah I know) and it came up with "With our next destination set amidst the asteroid fields, the vastness of space becomes our canvas, ready for exploration" creating an allegory for our adventures in space as if we are artists "painting our adventures onto the canvas of reality". Idk but I thought it was cool.
What is the meaningful difference between something truly being learned, and mere "copying" of "tiny pieces of material" that have been observed in the past?
Could learning exist in a vacuum universe, with nothing to copy?
Is there any LLM output that could convince the author of originality, or will it never be convincing? Can the author always tell whether he's interacting with one?
Questions around the emergent properties, sentience and creativity of AI systems should rarely be discussed as black and whites.
I don't think ChatGPT is sentient, but is it more sentient than a rock? I mean, maybe? At least I would be wrong to say a hard no to that. Similarly questions around the ability of LLMs to reason or produce original work probably isn't black or white either. LLMs do seem to show some signs of basic reasoning ability and have shown they are able to produce unique works, but clearly there are limits to both.
Perhaps this isn't unique to AI though. A "bad" artist is arguably one that is too heavily influenced by another and lacks artistic creativity. But does that mean what they do has no creativity? I think not, but that often is the black and white language people will use.
It seems to me ChatGPT is neither stealing or doing something all that creative.
Remixing has been part of culture forever. The author is really just grousing about the fact that machines can remix now.
- "They call themselves an artist but they are not that good." - "Copying and pasting shell scripts in a terminal does not a software developer make." - "The story that writer created is yet another permutation of <insert tale as old as time>"
You see what I mean? There is a good probability that today alone a significant percentage of content you saw online was AI generated and you were non-the-wiser and thought nothing of it.
It sounds good but mandating this will be death of AI in the west. This is relatively unprecedented situation where the use of copyrighted works actually helps you build tools useful for doing work. Training on DeviantArt makes the AI better at Photoshop-esque tasks.
Any country that doesn't implement this restriction will immediately be able to produce smarter more useful AIs.
> we should make databases not liable for the storage of any copyrighted material
I think this would actually be allowed right now legally speaking. That's basically a library or Google Cache. The hypothetical database wouldn't expected to be super useful because you can't "perform" any of the works outside of fair-use cases but it's up in the air of running inferences on that data (Google snippets) or training AI is a performance.
Google drives traffic so at least you gain something as a creator. Also amp has been rather unpopular.
How would you even know which way I did it?
Are we gonna fall prey to the scifi trope of "let's all be racist to machines", if so then I wouldn't hold it against AGIs to fall prey to the scifi trope of "I am gonna b evil now, bye bye humans".
Also, humans can be, and often are, found liable for copyright infringement or for piracy depending on how they conduct themselves. If a human was to reproduce a copyrighted book word for word, that would consist of copyright infringement regardless of whether it was done by rote memory, by copy and paste, or assisted by a black box LLM. Even if a human paraphrases another work they can still be found guilty of plagiarism if the paraphrase is still overly similar to the original source material. A human can also be guilty of copyright infringement if they use a copyright work as source material in certain ways. If I steal a stock image without paying for a license and add it in my Photoshop collage, I might be found to have pirated or infringed on the original image creator's property.
LLMs are trained on copyright data and can often reproduce that copyright data. It's an open question how we regulate this.
I personally think it would be fair for an artist or author to say their work was not licensed to be used in training a neural net or otherwise request to opt out.
Yeah. Nobody's talking about word-for-word duplication here.
> If a human was to reproduce a copyrighted book word for word...
Again?
> Even if a human paraphrases another work they can still be found guilty of plagiarism if the paraphrase is still overly similar to the original source material.
Go look up the dictionary definition of plagiarism. Notice the most crucial element, which you seem to have omitted here, and also notice that it's irrelevant to AI systems, which overtly acknowledge that they exist to generate derivative works.
> If I steal a stock image without paying for a license
Here's another version of your "word-for-word" analogy, which nobody else is talking about.
> I personally think it would be fair for an artist or author to say their work was not licensed to be used in training a neural net or otherwise request to opt out.
I am genuinely curious: how do you propose to enforce this?
why isn't existing copyright protection sufficient to regulate this? Photoshop can be used today to reproduce copyrighted data just as well.
This has nothing to do with "license to train". I do not believe existing copyright holders have this right granted to them by law - it is a right that is given to society for all works.
An artist learning a style, and producing another piece in the same style, is allowed today. This should be allowed, regardless of whether it is done via using an AI, or via years of training.
Consider an inevitable AGI with autonomy and no ties to a corporate. Can it not learn and write text whether in conversation, creatively or academically? If it's held to the same copyright laws that humans are then of course it's fine, in my mind. Hell, AI will be _better_ at avoiding infringing on others as they can store so much knowledge and process new knowledge (searching the internet for similar works) than humans are at infringing.
If this is still a problem then doesn't this just boil down to racism against the machine?
The author's main point seems to be that bits are colored[1], and that the process of LLM training doesn't "bleach models". In other words, ownership persists through training in a way that generated text mixes ownership rather than creating a new, ownable, artifact.
They also seem to have (extrapolating the theme) a concept of "smelly" bits. "They can also communicate with a person in a way that resembles actual conversation." ie: words have a "conscious"[smelly] category imbued upon them if they were generated by human or "unconscious"[unscented] if computer generated them.
This evokes a sense that "while they use words and form sentences like conscious people, they can't hold a conversation." Furthermore, this seems to be stated as an ontological fact, rather than say a derived and grounded in behavior such as, "They aren't currently capable of holding conversation because the output distribution doesn't match output distributions of human conversations."
These two views, that models don't bleach but they preserve smell, are fundamentally incompatible and diametrically opposed. How is it that color[ownership] persists through model training, but smell[consciousness] does not? We don't know - the author doesn't say. Jim just states both as fact without grounding either in rationale or justification.
Sure, we can take them as axiomatic, but then what's the point of the article then? There's a reoccurring disappointment in property/consciousness articles from tech-types who write on ownership and consciousness really aren't adding anything creative to the conversation. This is an extremely basic critique of the piece that the author should have thought through before hitting the publish button.
The second is if you include all of her works but also throw in every other book written in the last several years and ask it to create a work in the style of her franchise.
Is the latter somehow "better" than the former? I'd personally consider the former to be a clear case of copyright violation while with the latter we don't really have a way of knowing how much of her previous works were included. I'd still lean towards the latter being a copyright violation, and it's where I think we are currently with services like ChatGPT.
In that case, is that difference here that it's not humanely possible to train on world's information compared to a single human getting inspired by specific works. And hence an algo doing that is violation?
Quite possibly, it depends on what specifically is meant by style and the degree to which the new work is transformative.
Sharing a few surface levels elements like that is fine, but if a one page description could equally well be used to describe both works you’re well over the line.
lots of disney animations could've been described this way. And it's because they took storylines from well known folk tales and adapted it.
the idea of a teen wizard/witch going to a special school, fighting a villain, etc, doesn't seem to be copyrightable. While the russian ripoff version seems very similar and skirt the line between a copy and an original work, there's plenty of room where such a story could've been told and not be an infringement.
https://www.amazon.com/Tanya-Grotter-i-pensne-Noya/dp/569984...
You can’t copyright the idea of a wizardry school for preteens. You can’t copyright a style.
You can definitely trademark your own name.
A wizardry school for preteens isn’t on its own particularly unusual. But the question becomes how similar would the works be if the original was never published. You aren’t in the clear of a few paragraph book summery would apply equally well to both works. Barring the normal exceptions, Spaceballs is making reference to other works not just imitating them.
It’s also not a copyright violation because the question of how similar it is to the idea of Harry Potter is irrelevant with regards to copyright.
If it was too similar and it was actively confusing consumers then it would be an issue of trademark.
By making the protagonist a girl without glasses it does enough to differentiate itself that no reasonable confusion would ensue.
As long as the books don’t contain verbatim copies of text from Harry Potter they are non-infringing!
“Despite its reputation in Russia and the many books it has spawned, the series is not available in English translation, because of the first book having been judged a breach of copyright.” https://en.wikipedia.org/wiki/Tanya_Grotter
So, respectfully you are simply mistaken.
That doesn’t mean I am incorrect about style not being copyrightable.
Anyway, many Russian works would be considered copyright infringement in other countries. The country however has minimal interest in respecting foreign copyrights. So if that’s your benchmark I can see why you might assume more leeway than actually exists.
Otherwise I would have done the bare minimum of reading the Wikipedia entry on the book. All I did was Google for “Harry Potter knockoff”, saw that the book was on Amazon, and thought that was enough.
Reading the Wikipedia article it is clear that this series took much more than just the idea of a wizardry school and copied verbatim plot elements. It was not the right example to use.
Here’s what I mean by style. Reggae music. How different is one reggae song from another? What about when a reggae musician sells the rights to one song and then write another song in their own style? Are they infringing on the previous work they just sold to a music publisher?
Here’s what happened when John Fogerty was sued for sounding like John Fogerty:
https://www.mentalfloss.com/article/27501/time-john-fogerty-...
This logic seemed pretty sound to the jury. It only took two hours of deliberation for the jury to determine that the two songs didn’t meet the legal standard of being “substantially similar” that would have constituted copyright infringement. The Fogerty camp let out a collective “huzzah!”
The case was litigated rather than being thrown out because style is legally recognized as falling under copyright. “In 1993 the United States Court of Appeals for the Ninth Circuit shot down that appeal, though, on the same grounds—the original suit had been neither frivolous nor brought in bad faith.”
The Supreme Court didn’t disagree, they ruled on a different matter.
That Ms. Rowling and her publisher won doesn’t mean that style is covered by copyright either. It could just mean they had a more expensive legal team.
One of the key elements of copyright is this doctrine:
https://en.wikipedia.org/wiki/Idea%E2%80%93expression_distin...
An adventure novel provides an illustration of the concept. Copyright may subsist in the work as a whole, in the particular story or characters involved, or in any artwork contained in the book, but generally not in the idea or genre of the story.
Also, let’s take a step back. What do you think is a fair outcome, that once John Fogerty sold the rights to one CCR song that he is no longer able to write music? Or that he needed to learn to write and play an entirely new genre?
Does the first music publisher to buy a reggae song now own the rights to every single reggae song produced after?
What are you arguing for?
> Or that he needed to learn to write and play an entirely new genre?
As the court case you mentioned shows genre isn’t specific enough to be problematic. I can’t draw a clear line in the sand and say this is safe because ultimately that’s up to the courts to determine in each instance. But, it is important to understand the general areas that are risky.
https://en.wikipedia.org/wiki/Pharrell_Williams_v._Bridgepor...
Appellate Court Judge Jacqueline Nguyen wrote a dissenting opinion: “that the judgement allows for protection over musical style”
That doesn’t mean “style” is enough to lose a copyright case, but shows such things are potentially problematic.
Here's a concise explanation of why this was a terrible ruling: https://www.youtube.com/watch?v=-1COYitP8hI
If the entire music industry operated in such a manner it would suffocate creativity.
I think what you’re really trying to do is win an argument that supports your opinions on LLMs.
Generally speaking, the Robin Thicke case was an aberration. If it was the norm then the music industry would look incredibly different. As in, there wouldn’t be a music industry.
We can go back and forth cherry-picking court cases or we can discuss the doctrines that courts are encouraged to follow, like the idea-expression distinction.
An overwhelming number of copyright cases have been decided based on this doctrine! It is disingenuous to suggest otherwise!
Winning a court case isn’t safe, getting the court case thrown out before it needs to be litigated is. You really don’t want to be in a position where someone with deep pockets can make a reasonable argument for infringement.
The question of LLM’s is really secondary. It’s likely they could win, but that in itself is very dangerous territory.
What about a drum pattern like a four-on-the-floor disco beat? What about entire genres of electronic dance music?
What about chord progressions? Should songwriters avoid the I-IV-V? Do you know how many songs are ii-vi-V-I?
What about literally any classic country song?
Do The Strokes owe the band Television royalties for imitating their driving 8th note, interwoven guitar riffs?
Here's a good conversation around the roots of an aphorism typically attributed to Picasso, "good artists copy, great artists steal.":
https://quoteinvestigator.com/2013/03/06/artists-steal/
Are you a songwriter or a musician? If so, how can you make music without imitating any number of factors? Do you make your own instruments? Do you have your own harmonic system with your own scales?
I can't think of any creed that is more destined to result in terrible art than "playing it safe"!
EDIT: Oh, "don't imitate anyone" is a close second! The one-two punch of playing it safe and not imitating anyone is just, I mean, I can't even...
The line isn’t anywhere close to the hyperbole you and many other people are bringing up. Don’t copy other peoples work isn’t some unrealistic burden, it’s the basic foundation of copyright.
I am a songwriter and a recording artist. I and every other musician I’ve ever worked with have conversed in the language of our predecessors. We constantly listen to other records while making our own and borrow drum patterns, guitar sounds, chord progressions… just as all of our favorite musicians did before us.
If you don’t do this and you go off on some quest for pure originality, you are engaged in pure narcissism and you will make terrible art.
Show me your favorite record and I can break down where every single element of every single song came from.
I'm going to guess that you have basically no experience making music. Here's just a couple of the records I've made:
https://williamcotton.bandcamp.com/album/melencolia-i
https://willieandallie.bandcamp.com/album/between-a-rock-and...
However, your actual defense is a lack of serious earnings. If you habitually copy significant portions of your work don’t assume it’s actually legal rather than simply not worth suing over.
If you’re concerned about actual credentials my connection to this stuff is through major motion pictures which means serious paranoia around IP. They are actually concerned about being sued and take actual precautions.
Have a listen to my albums and you can judge for yourself how much copyright infringement I’m guilt of…
If we had these LLMs 30 years ago would J.K Rowling's books ever have existed?
All your question does is give people an excuse to make up alternate history in order to argue whichever side of the argument they already believe. It serves little purpose in the dialogue of defining AI and human creativity.
If you feed a 'young author' the works of a famous current author, E.G. Rowling, and ask them to write a new work in the style of that reference author...
Or if you include the current N 'Best Sellers' and ask them to write whatever's popular.
I'm not so convinced creativity is any more magical than a naturally evolved and refined version of what computers are starting to approach. Humans naturally add far more background, results filtering, and implicit selection parameters (personal biases / preferences).
Maybe the correct line is somewhere around; tracing (direct copying) is bad, but freehand (from memory) is OK. However computers inherently have more perfect memory than humans; how precise is the detail from the source work? Is there a meaningful threshold where the memory of a work has decayed from a representation of the source work?
Humans learn by "theft". We learn language by observing word patterns created by others (books, speeches, TV shows) and correlating them to outcomes. While ChatGPT doesn't yet display human-like intelligence, it's clear to me that it is taking a step in the direction of how the human brain actually learns.
> "How do I make a peach cobbler?"
"To make a peach cobbler, run your fingers aggressively over the bosom of the graham cracker crust...thrust the pie pan deep into the oven; don't worry about spilling the mix."
However I feel like we are at the verge of a tragedy to come
Back in the day, broadcast radio stations had everyone believing that we shouldn't have to pay for music.
In 2023, we have youtube to thank for that.
> If I own something, why shouldn’t I be able to share it with whoever I want?
Good argument, especially since many intangible goods, such as MP3 files, are non-rivalrous and non-excludable.
so the deeper question is 'what does it mean to own something'.
For a physical good, it's fairly straight forward.
For a digital good that can be copied infinitely and perfectly, it's not. The current definition of ownership seems to be broken down into a set of "rights". Sharing it with somebody else is one such right.
So when you purchase an mp3 off a store, you're purchasing not the ownership, but the rights to listen to it. More importantly, you're not purchasing the rights to share it with somebody.
But the idea of the song - whatever that might be - is not a right given to the owner to be sold (at least not currently). Therefore, someone who purchased the right to listen to the song _should_ be allowed to learn the idea of the song, and reproduce the idea in another form.
As I understand it, the Audio Home Recording act legalized mix tapes and mixes on DAT and "music" CD-R media. It's why "music" CD-Rs cost more (the DART royalty.)
In fact Apple still provides software (Music/iTunes) that allows you to burn purchased songs from the iTunes Store onto CDs:
https://support.apple.com/guide/music/intro-to-burning-cds-a...
Initially they hadplaced limits on how many times you could burn the same song, but Apple removed the restrictions and eventually upgraded all of iTunes to DRM-free (but watermarked) iTunes Plus in 2009.
Unfortunately it does not support burning "rented" songs from Apple Music or Spotify. I guess that's what Audio Hijack is for.
And sadly iTunes never got an upgrade to a lossless format while Apple Music streaming did. I guess that's a good reason to buy tracks from other sources.
I also object to its clickbaity format of "[Contentions thing A] is not [completely unproven hyperbolic thing B]; it is [completely unproven hyperbolic thing C]".
They were also successful in brainwashing people that this is a good thing. Phrases like everyone should get paid for their hard work are thrown around. And the brainwashed masses are frustrated to realize that they get only pennies for their art. They are embittered that they can't stop others using their art. They keep shouting that word: Theft.
Art is fundamentally derivative. It is my deep conviction that nobody has the right for a monopoly on usage, even that group, even if law says otherwise.
But hard work deserves compensation! I hear the cry.
Look at the more famous Youtube producers. They get compensation but at which price? They don't work for themselves anymore. While creating stuff they think about how to reach even more people. They lost the truth. Instead of truth they say, don't forget to click the bell.
Or look at the people working for free like parents caring for their children. They cook, clean, educate not for a lot of material gains, even worse, they have people telling them that what they do is not real work.
There are other examples of hard working people not getting compensation nor recognition or not a lot.
Or down the road a bit, an AI trained on and able to draw conclusions from all our scientific papers.