LLM-based music generators do what musicians do - listen to a lot of stuff and then generate their own music. Most popular music isn't that original. It's mostly about branding.
LLM-based music generators do what musicians do - listen to a lot of stuff and then generate their own music. Most popular music isn't that original. It's mostly about branding.
I love this analogy, because it looks so benign, while being so off.
Question: How many hours of music an LLM can listen and ingest in an hour?
Answer: Probably 10x more than a human, and will ingest it way better in one go than a human who needs 5-10x concentrated listens to completely understand what's going on at all layers.
So, in practice, an LLM can ingest 50-100x more music in unit time, when compared to a human.
"An LLM is like a musician. Just listens". Indeed, that's true, if I wince really hard, so my head hurts.
BTW, I'm all for respecting copyright, and no, AI training is not fair use, because AI is not a human who consumes reasonable amounts of information and understands and derives it in a unit time.
Please illustrate why "LLMs ingest training data faster than humans and have better recall" has anything to do with fair use.
Notwithstanding the provisions of sections 17 U.S.C. § 106 and 17 U.S.C. § 106A, the fair use of a copyrighted work, including such use by reproduction in copies or phonorecords or by any other means specified by that section, for purposes such as criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research, is not an infringement of copyright. In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include:
1. the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
2. the nature of the copyrighted work;
3. the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
4. the effect of the use upon the potential market for or value of the copyrighted work.
The fact that a work is unpublished shall not itself bar a finding of fair use if such finding is made upon consideration of all the above factors---
Let's polish over the "scholarship, or research" clause for a moment, because none of the commercial entities are doing it for research. Assuming the fair use clause didn't fail there, let's focus on two important points:
3. the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
4. the effect of the use upon the potential market for or value of the copyrighted work.
First of all, an LLM remembers almost all of its training material verbatim, or almost verbatim, which is undoubtedly a substantial amount of the work ingested*, and secondly, the effect of such systems on the whole market or job ecosystems is disruptive. You can create something in the style of someone from your couch for free or $10/mo.Whatever you do, you can create derivative works of the ingested works, or derivative works containing ingested parts verbatim at great speed with this "Shiny LLM thing".
A human neither can ingest that amount of information, nor retain all of it that perfectly, hence limiting the inspiration from these works, and forces one to add their own original ideas or inspirations rooted in other sensory inputs (feelings, sight, own experiences, etc.), not reducing the value of the work used, or not damaging the creator of the work the creator got inspiration.
Incidentally, we have a word for cases this inspiration gets a bit too far: "plagiarism".
From that point, an LLM, which remembers everything verbatim and mixing them together is arguably making plagiarism from many sources to create their own works, while being unable to cite them accurately. A very useful property, indeed.
*: I know how an LLM stores its data in its weights, but even without the data itself, the weights can construct the training data verbatim in many cases, so it stores its training data. Not in its original form, but in a derivative and reversible form, not unlike compression.
[0]: https://en.wikipedia.org/wiki/Fair_use#U.S._fair_use_factors
I think this is a weak point in your argument. I support the rest of your reasoning, but this point can be attacked. For plagiarism to exist, the original work has to be easily detectable by a human. We've seen this sometimes be the case (see NYT lawsuit), but often times, the line is blurry (see Ed Sheeran lawsuit [0]). There's a reasonable argument to make that the combination of existing ideas by an AI creates something novel. It's what arguably what artists do, too. There's even a famous book called "Steal like an artist" [1]. As such, some output by AI can be plagiarism, but the very way modern AI works may not automatically mean everything is plagiarism.
The people that need to change this status quo - like it or loathe it - are the people that want to profit from spamming streaming services with zero-effort derivatives of popular music (people that do it for fun or use 'AI' tools for post production or loop-variation will continue to be fine). I have a lot less sympathy with them than the many artists who've lost court cases for 'inspiration', and if they're relying on the whole "models learn just like people" canard they're going to be in a weaker position than usual, because not only is that deemed insufficient for people, but it's unambiguously untrue that machines "listen" like humans do (they're bulk processing files, not experiencing consonance or dissonance in the ear) or that they "perform" as opposed to stochastically transferring elements of style. Even the most cliched blues going has human reasons for note choice and intonation other than "weighted average of other blues with overlapping description"
From what I've heard about human cognition, we come up with a word-based justification after we decide what to do, it just that the research involved a brain scanner to find this out because it doesn't (usually[0]) feel like this is what's going on on the inside.
So the real reason may well be "weighted average of other blues with overlapping description", followed (not preceded!) by a language model producing some convincing words that can communicate the mental state to other humans.
IMO, the main difference between AI and humans right now, is that humans can do human things with far fewer examples. AI has the speed advantage that lets them read all the words in the world, listen to all the music ever recorder, watch all the videos ever filmed… but are so stupid[1] that they need to, too.
[0] I do sometimes notice this in my own brain: I can have a fully-formed sentence before my inner dialogue tries to (for a lack of a better word for the internal voice) "speak" it — if I then try to skip the inner voice because I've already had the thought and turning into an imagined voice is clearly a waste of time, the part of me which is the inner voice is, for lack of a better word, frustrated.
[1] Stupid on this particular axis. Intelligence isn't fantastically well defined even in humans, is badly defined cross-species, and doesn't have a singular definition at all for machines, so it is perfectly valid to consider intelligence to be a question of "what can it do?" rather than "how fast does it learn?", for which many current AI (there's no G, so narrow AI that only plays chess is fine) "smarter than humans".
I mean, if humans need to listen to far fewer examples to create music than an LLM needs to parse sound files, they're obviously not doing the same thing, are they?
Similarly, it's unambiguously false that a human blues is the weighted average of the sound files of other blues. There are structures which are essentially identical to those used by other artists for reasons of easy improvisation, there are selections of notes played over the top which are a function of fundamental musical properties of consonance and dissonance, scale runs the instrument is suited to as well as licks that may have been directly borrowed, but the final output is dependent heavily on the actual capabilities of the artist to play and singer, which obviously aren't simply lifted from their favourite artists' recordings. Even if it's highly derivative which it almost certainly is, it's not remotely the same progress as a matrix generating a sound file from a combination of other tagged sound files and the input "12 bar blues with Stevie Ray Vaughan guitar work and BB King vocals"
If I need much longer to run a marathon than an athlete, is that because my body is not doing the same things as the athlete, or just because it's doing all the same things but doing them worse?
If a (clinical) moron tries to learn a song, and it takes a lifetime, why? (Do we even understand human brains well enough to answer that? Not my field).
> there are selections of notes played over the top which are a function of fundamental musical properties of consonance and dissonance, scale runs the instrument is suited to as well as licks that may have been directly borrowed
How do you learn what these things are? What do "consonance and dissonance" mean in terms of why they make you feel the way they do, and why can't an AI learn those things from examples? The qualia, that AI hopefully doesn't have but we can't tell because we don't have a test for that?
Danse Macabre (Camille Saint-Saëns) was rejected at the time for how it sounded. The early cultural rejection of Blues seems to me like it was more probably racist excuses than artistic preferences, which is why I'm picking a non-Blues example.
> reasons of easy improvisation … [all the other things in the previous quote, I hope this is a fair cut] … the final output is dependent heavily on the actual capabilities of the artist to play and singer, which obviously aren't simply lifted from their favourite artists' recordings.
If I put the AI into a robot (or simulated) body, and the limits of that body constrained the capabilities of the output, those limit themselves make it OK? Constrain it to MIDI rather than a full waveform?
This doesn't seem like it would change the answer, to me.
If it's because you have to generate the concept of legs from observing other people run marathons before emitting a digital facismile of a race, I think we can be confident you're not doing the same thing as the other athletes.
How do you learn what these things are? What do "consonance and dissonance" mean in terms of why they make you feel the way they do, and why can't an AI learn those things from examples? The qualia, that AI hopefully doesn't have but we can't tell because we don't have a test for that?
Consonance and dissonance are physical properties of sound resonating in our ears, our emotional reactions involve observable physiological reactions like dopamine release and I think we can confidently reject the premise that a static model of relationships between wav files which instantaneously generates another wav file is doing the same thing
Nobody is arguing musical fashions don't exist, because they do, they're arguing that AI unambiguously does not "listen" to music and a human performer unambiguously isn't even capable of outputting a mathematical translation of all the sound files they've absorbed, never mind solely doing that like a large model trained on other people's music [and text]
> Consonance and dissonance are physical properties of sound resonating in our ears,
If a duff chord is played out of hours in a full morgue where no living person can hear it, is it still dissonant?
> our emotional reactions involve observable physiological reactions like dopamine release and I think we can confidently reject the premise that a static model of relationships between wav files which instantaneously generates another wav file is doing the same thing
I'm not so sure.
Observable physiological reactions are the easy part, but they don't solve the "hard problem of consciousness"[0]. Various chemicals, including dopamine, modulate various neural activities and let our brains focus on things and learn in certain ways. Attention heads allow a big confusing pile of linear algebra to "focus" on particular things and learn in certain ways — what's that like on the inside? Is there even an inside for it to be like something? Nobody knows. I've only heard one suggestion for how to test it that isn't immediately obviously flawed: https://www.youtube.com/watch?v=LWf3szP6L80
(I don't know if Ilya Sutskever came up with this test, but he's who I first heard promoting it).
We're made of atoms. The chemistry of life led to us having feelings as the result of evolution, a process that has no intentionality. An AI may, or may not, reproduce that. In the absence of a test (without which we can't even determine the presence or absence of qualia in mammals with brains of similar complexity as SotA AI models), I'd be just as skeptical of anyone who says some AI does have it as anyone who says that AI doesn't have it. At the same time, given how we got it, I think our human brain chemistry is far less magical and more "mess that works" than our egos want to believe.
I'm happy to call current AI "fast but stupid". Anything else, including future AI, I'm not so sure about.
[0] https://en.wikipedia.org/wiki/Hard_problem_of_consciousness
Obviously, since we can't prove the morgue isn't consciously listening to it... :p
We can't prove that Napster or a counterfeit CD isn't conscious to everyone's satisfaction either, but fortunately for copyright-holders, that isn't actually how the burden of proof on a "it's just like humans learning how to perform" defence works. Legal systems deny that things that probably are conscious are human-equivalent all the time anyway, that's why you can eat or sell a cow but you can't sue it or assign copyright to it.
But even if the burden of proof was reversed, an argument that an NN consisting of static values stored on disk that outputs a discrete and (similarly) static sound file on demand before returning to being static without even a state update isn't the same thing as a human mind/body constituting long-running chemical processes performing music in continuous time seems.... pretty sound.
People do things with LLMs. People write and run code that creates LLMs.
These words "train", "learn" are ML terms that attempt to explain a concept using an analogy. The LLM isn't a person.
Speaking of the LLM as if it is a person, has the rights of a person, or should be legally treated as a person, is a category error.
Whens the last time you saw a human generate a photo realistic image in bits and bytes? Struggle to find the words to express a loss?
I love your comment.
We have plenty of content with attribution to train such a system - books, art, scientific papers, media articles, even social network comments, we know author and date. We can index everything by author and search against the index to find influences. We can train "text -> author" and the use it for "new text -> attributions".
But will such a system ultimately enable or block creativity? Even past works stand to be revealed in all their influences. Might be embarrassing. On the other hand generative models will be safe as they can sample around problematic attributions and still complete their tasks. We might end up having to use AI for writing in order to ensure copyright safety. The irony.
"I heard you don't like AI in creative fields so from now on you need to use AI to ensure other creative's ideas are not used by mistake, or risk getting sued."
Maybe it is a lost fight and the value has to be found somewhere else though. Anything physical like deployment, hosting, that AI cannot do, at least for the foreseeable future.