My way of "giving this serious attention" is through pre-registered, falsifiable, repeatable, experimentation, which anyone can look up on osf.io because I use my real name. I'll bet you that non of the randos in this thread do as much.
To all of the randos: unless you have data... it is just an opinion.
Glib as well, but this one hits home a lot harder. Well said.
The problem is that people's words are MUCH more predictable then they would like to believe. And that truth upsets them.
In addition to having created models, I also write books and articles. Probably more than most people commenting here. I have a firm grip on what actual copyright law is and the pros and the cons of it.
> The problem is that people's words are MUCH more predictable then they would like to believe. And that truth upsets them.
I'm not offended. I do think it's a little weird that you seem to think "training on a bunch of stuff that includes a set of words" and then "predicting" those words exactly is somehow okay because theoretically it might be extrapolating the exact same words from combining other ones. I'd argue that if a model trains on data, and then reproduces exactly a large subset of that data, the bar should be pretty high to prove that it's not copying, and "you don't understand because you didn't implement this" is not a good basis for law.
> In addition to having created models, I also write books and articles. Probably more than most people commenting here. I have a firm grip on what actual copyright law is and the pros and the cons of it.
I'm not convinced you have a firm grip on the idea that no matter how smart you may be, "just trust me bro" is a pretty terrible strategy if you're actually intending to convince anyone of anything. If that's not what your goal is here, it's not clear why it's worth your time to respond to other people's comments when you clearly have so many other productive ways to spend your time.
I am asserting it is Charles Fort's "steam engine time". Far from a crank position. It is one that bears serious consideration.
I think it points to an interesting trend either way. People are less tolerant of machines. Failures of machines are reviled because of their nature, even when the overall problem compared to humans is less. For example, self driving cars. If self driving cars halve traffic deaths from reckless driving but it occasionally mows over a family of four in broad daylight for no apparent reason, society will overwhelmingly reject the technology.
Basically, I dont think people will ever be satisfied even if we prove "its just doing the same thing we are." It's going to be held to a higher standard.
I don't think we should "get over" the fact that modern SOTA models couldn't exist without being trained on protected works.
That someone, at some point, paid for.
I'd like to understand why I can't use a song in one of my videos without permission/payment, but an AI company can train models using that song without having either.
I'm not anti-AI. I'd just like to see companies play by the rules everyone else has to follow.
Because training isn't redistribution.
You can also listen to the song and make a new one that sounds similar, just like the AI can.
Answer: They did not. That is literally why there are dozens of ongoing lawsuits in progress.
Because when you say you are “using” the song, what you mean is that you are distributing copies of the song, which is protected by copyright.
When AI companies train on the song, the model is learning from it. Outside of the rare cases of memorisation, this is not distributing copies and so copyright doesn’t have any say in the matter.
Learning isn’t copying, so copyright doesn’t get involved at all.
The New York Times is suing both OpenAI and Microsoft for copyright infringement. The Authors Guild is suing OpenAI. Getty Images is suing Stability AI. Disney is suing Midjourney. Universal Music Group and Sony have filed suits against multiple AI companies.
> so copyright doesn’t get involved at all.
The dozens of ongoing cases that discredit that statement.
Your objection doesn’t make sense. In the event that an AI company loses a lawsuit for copyright infringement based on simply training on copyrighted works, the answer to you saying you’d like to understand why they can do it and you can’t is simply “your premise is wrong; neither of you can”.
I object to your statement that "copyright doesn’t get involved at all" when that is objectively untrue. If that was true, many of the world's largest companies wouldn't be spending tens of millions of dollars to have that question answered in court. Go to any law-focused forum, and you will find attorneys arguing over these questions.
To train a model using a book, you must first obtain a copy of that book. Did OpenAI purchase a copy of every book not already in the public domain used during training? They did not.
Some of the suits I mentioned claim that OpenAI literally stole copies of books to train its models.
My point is that the copyright question has not been answered. If the NYT, et. al. win, it will be a watershed moment for how AI companies pay for training data moving forward.
You're right, it's an unjust situation. And you may note that no one else besides the AI companies has made any progress at all towards changing it.
Copyright will soon die, having outlived its usefulness to society. Whether the knife is held by someone named Stallman or someone named Altman is of little consequence.
[0] https://archive.org/details/hisyo00simo/page/n1/mode/2up