Apple trained AI models on YouTube content without consent
9to5mac.com
9to5mac.com
Scraping for LLMs should be illegitimate, and self-serving analogies to human learning should be rejected. Whatever precedent that allows for search should definitely not apply to LLMs, because LLMs have different characteristics: the traditional search is more like movie review, it borrows from the movie but in the process directs you to the creator's business (where he can profit); while the LLM is like a pirated copy of the movie, whose point to keep you away from the creator's business (so someone else can profit). At their core LLMs are exploitative in a way search is not.
> Is there a real legal precedence that has set a clear line here?
When circumstances change, the law can change to achieve a just result.
Software engineers seem to like to mis-imagine the law like physics: unchanging, and if you can figure out some loophole to achieve something you get to do it forever. But that's wrong, and the law can and should change.
Not sure how this is self-serving or a reference to human learning. This is a pretty emotional reply and it makes it hard to see the points you are trying to make. I think there is a discussion about the legitimacy of scraping in general that is ongoing, which is a lot of what I brought up in my post. Training using scraped information should be part of that discussion.
An argument I frequently see made to defend companies who take all kinds of people's content to build LLMs (which then compete with the people's whose stuff they took) is this: 1. People learn from seeing examples, and that's protected. 2. Machine/LLMs learn from the content (the trick being pulled is to equate human learning with machine learning). 3. Therefore OpenAI, Google, etc. should be able take everyone's stuff to build the tech they want to sell / I want to use / fulfills my nerd fantasies.
It's self-serving because they don't care about the harm done to build the thing they want, so are trying to excuse that harm away.
> This is a pretty emotional reply
What's wrong with emotion?
"It’s important to emphasize here that Apple didn’t download the data itself, but this was instead performed by EleutherAI. It is this organization which appears to have broken YouTube’s terms and conditions. All the same, while Apple and the other companies named likely used a publicly-available dataset in good faith, it’s a good illustration of the legal minefield created by scraping the web to train AI systems"