Apple, Nvidia, Anthropic Used Swiped YouTube Videos to Train AI
proofnews.org
proofnews.org
so basically "we stole from a thief therefore we didn't steal" excuse?
Don't humans basically do the same thing when attempting to create new music — they derive a lifetime of inspiration from the works of others?
Maybe go after the application, not the technology? Someone uses AI to explicitly plagiarize an artist’s content? Sure, go ahead & sue! But limiting the growth potential of a whole class of technology seems like a bad idea, a really bad idea actually if your military enemy had made that same technology a top priority for the next years …
If I train people to draw anime from a book on how to draw anime, and ask them to start drawing work related to Bleach (e.g.), have I disrupted the market for the original works of Bleach?
Is it that disgusting to you to discuss the law that you want to derail it by talking about how sad it is to ask?
The question is where on that spectrum does current AI training lie, and where is the cutoff between fair-use and unauthorized commercial use of copywritten works.
Today's AIs are not the same as a human creating original work. Even humans have to be careful to not reproduce existing works too closely or they also get blamed for plagiarism.
Are we normally required to label sources when referencing other copyrighted materials, whether in songs or movies or otherwise?
As we know with scraping cases, the amount data and time also may play a role in determining fair use (think in terms of buffet ettiquite. "all you can eat" does not in fact mean "you can eat it all by yourself"). Funnily enough LinkedIn (owned by Microsoft) did argue successfully in court against scraping a website.
If I take ten of your copywritten photographs and stack them on top of each other in Photoshop with transparency, the output is not a simple reproduction of your work. If I sold that for commercial purposes you would be upset with me and likely have a copyright case.
That's an obvious example, but my point is there aren't super clean-cut definitions for these things, and it's not settled case law yet which side current AI training and content generation falls under.
Most famously might be Andy Warhol's Campbell's Soup cans. You can find plenty more though. Product labels, magazine covers, pasted together as "art" even thought each part is copywritten.
If the question is about what is "fair" I don't see how you could be surprised that artists, journalists, musicians, Youtube creators, would object to huge tech companies using their stuff without permission to replace them. It is entirely to be expected that many people find this unfair.
If enough people file this to unfair and outragous, even if the courts found the "fair use" arguement cogent, the laws could be changed.
Every music service has something like this, are they delivering just the value of the streamed music? Great then they only owe those royalties. Are they delivering the value of EVERY song they trained on every time a new song is chosen? I sure wasn’t asking that question until generative results became the product.
... which they presumably listened with permission
Therefore isn't training AI on this basically poisoning your own model? The caption quality is good but there are mistakes in pretty much every video I watch with captions.
https://github.com/Dicklesworthstone/bulk_transcribe_youtube...
E.g.: Nike needs to produce a large amount of clothes. They hire an oversea company who commits to the order. They set strict rules -- no child labor, certain quality controls, etc... This company then subcontracts anyway possible and delivers the order, gets paid, and dissolves. Messy but Nike's hands are clean.
With AI, same thing but with videos and other forms of data.
Hence why a question "did you train with Youtube?" to a certain CTO is so difficult to answer.
Make it so they're automatically guilty if they can't provide a definitive negative answer.
And yes, regardless of results I agree there should be new laws made. But we know Congress in the US this year has been a roller coaster, to put it lightly. And I don't even think this is top 5 of what congress needed to codify into law properly. So all the short term work will be the judicial branch interpreting what few laws we do have.
What enhancements are you thinking of?
IMO data is Google's biggest moat in the AI race - and I suspect they'll do whatever they can to keep it.
Meta is the one with the huge data moat.
Personally been using fabric ai tool, since it can summarize youtube videos, so I dont have to watch an hour+ video or read a very long article/journal, just gives me a summary, top talking points or even break it down for tech points.
Citation please.
One could say the quantity. We're currently dealing with statistical learning models that require a huge quantity of training data. This is temporary. At some point you will be able to train an ML system with less because humans can be trained with less. What then?
For example, Jacksepticeye is listed as having their videos used. Looking at the channel, it seems like a lot of it is is recordings of them playing video games.
Is the company that produced these games being compensated?
As much as I hate react videos, at least the situation is the same. You know where the source is from and can go to the original if you wish.
Show me the attribution on your generated content for any of these creators.
IANAL, but recording portions of copyrighted content or using excerpts thereof is covered under fair use.
It is not yet known whether reproducing copyrighted content in substance or style using generative AI is covered under fair use.
There was a period of time several years ago when some publishers were not allowing some kinds of streaming, and there was a need for them to post public statements allowing things like Let's Play videos. Even now you'll have the occasional game where the publisher just says streaming isn't allowed, or is only allowed under restrictive terms.
Given the relative simplicity of video game worlds, it should be far easier to generate those than than photorealistic video (i.e. Deep mind Veo, OpenAI Sora).
Yes, it might just saturate the world with low quality content, leaving the good stuff still distinguishable, but many content business models are built on low quality content.
IMO, the most problematic part is generating content in the style and voice of the original human creator of the copyrighted content.
Unfortunately, this is also currently among the most user-attractive for generative AI trained on copyrighted content.
Most articles will link to sources from others and then build on top of them for their own article. Actually giving sources for your work. That isn't swiping the content to write an article. Yes there are bad actors in that regard, but most play by the accepted rules here.
With the AI work, attribution of source is gone. Who did the work that you are benefiting from is gone. The people that benefit are those that made the AI and the people using it, skipping the source for the training data.
I'm not going to grok through all their thousands of videos, but these specific creators sure didn't get this big on reactions only.
& a thumbnail with the logo of Apple at the top-right and their face doing a weird expression, occupying 75% of the .png with a single bright color background