OpenAI transcribed over a million hours of YouTube videos to train GPT-4
theverge.com
theverge.com
If push comes to shove, openai can acquire all the data they need with microsoft money.
But the little guys can't. So if we let them get away with it, we are really making sure the little guys can compete!
> Google also gathered transcripts from YouTube, according to the Times’ sources. Bryant said that the company has trained its models “on some YouTube content, in accordance with our agreements with YouTube creators.”
> The Times writes that Google’s legal department asked the company’s privacy team to tweak its policy language to expand what it could do with consumer data, such as its office tools like Google Docs. The new policy was reportedly intentionally released on July 1st to take advantage of the distraction of the Independence Day holiday weekend.
There needs to be legislation to reign in the practices of big tech and also this increasingly common dark pattern of forcing updated terms on customers. They can get away with this due to a lack of feasible competition and their massive capital. Only regulation or taxation might resolve this issue so that control lies with content creators.
Two perspectives: (1) because google is paying for the storage, processing and delivery of these videos, so it makes sense that they could do something with them that others can. And (2) because maybe you uploaded those video to YouTube specifically because you find Sundar Pichai vision extremely interesting so you definitely want to be on his platform and you find Sam Altman so obnoxious so you don't want him to use your content.
If the USA and Europe take this route, what happens next? It's conceivable that China might not follow suit, opting instead to leverage all available online data to boost its AI development. This could lead to a scenario where China manages to develop a super AI, while Europe and the USA get caught in a complex maze of licenses and expensive, specialized micro-AIs.
Faced with an advanced and accessible Chinese AI, what will users around the world do? It's plausible that despite initial concerns, many will end up using the Chinese AI, thereby giving it even more power, data, and growth capacity.
This brings us to the importance of balanced regulation. The European Union, for instance, is often viewed as having hindered innovation with excessive regulations over the past decade. If the USA were to follow this path, they might risk limiting their own capacity for innovation.
Excessive regulation thus emerges as a double-edged sword: on one hand, it aims to protect privacy and user rights; on the other hand, it risks stifling innovation and making one's companies and technologies less competitive globally. In this scenario, the question becomes: How can the USA and Europe balance data protection needs with the necessity to remain competitive in the global AI arena?
Moreover, this discussion isn't just about licensing MP3s; it's about managing common knowledge—that which is accessible to everyone, including competitors in different jurisdictions—for the creation of an industrial revolution. Before clamping down with regulations, careful consideration is needed. Licensing mechanisms can be introduced later if necessary, but they should not provide an undue advantage to competitors. Balance is essential.
You have written "The European Union, for instance, is often viewed as having hindered innovation with excessive regulations over the past decade" - but how do we know if that's true? Maybe its regulation didn't do much and the US growth is simply caused by the printing machine for the global currency combined with an extremely homogeneous market (all ppl speak the same language and globally a lot more ppl speak english compared to say german/french) and it's part of one country compared to EU that's a group of countries operating under some agreements but with still different laws/cultures. How do we know that specifically EU regulation is the root cause of 'less' EU innovation compared to the US and not the other things?
that tracking system is not an industrial revolution ...
you are comparing different thing ( that was my mp3 licencing reference , the compromise on preventive legislative micro-management must be evaluated based on importance )
states choose what to regulate based on the need for control/growth/safeguarding/... everything affects different points differently,
make comparisons with others laws or regulations on other topics must take into consideration the implications that these different categories have on the compared subject
so "they" can have their AI that makes new molecules, new medicines, new materials, increases productivity by 5x , in the meantime should we separate all the scientific research papers because otherwise they will devour us? , so we move backwards but are we giving $5 more to everyone who wrote a book that no one has read? so they can buy the chinese vpn to access their ai to write the next book for the next 5$ ?
What strikes me, why do you think that EU/US can't do all this stuff if they put in place some regulations (like molecules,medicines,materials, etc...)? Regulations (normally) should limit only a subset of usecases while allowing others, don't you think it's a normal approach instead of letting big corpos suck for free any public data(be it copyrighted or personal) and act unbounded while profiting from this? We know what corporate greed is, look at the insulin price in US vs EU, why don't you think that the same will happen with AI tools/corpos?
Great, why don't these companies train their models on that "common knowledge" then and leave copyrighted work alone? Would be a great option wouldn't it? Solves the whole problem in one go.
Personally, I think these companies are just stealing copyrighted works, and should be sued for that. It's against the law, it's pretty simple.
And the whole China argument... Sorry but it makes no sense. They could beat us in AI by doing illegal things, so we should do the same illegal things to not let them win?
China has a lot of slave labour too. By the same reasoning, shouldn't we introduce slavery to not let them win production?
Is it copyright infringement to count how many times each letter appears in a book?
I don't know that it is, and if it's not, then there is at least some line you can draw where mechanically reading and learning from a copyrighted work is not copyright infringement.
The question of whether training a transformer model is on the legal side of that line remains to be seen, but I don't think it's as clear cut as you make it out to be.
I see it as more similar, for example, to how when the first search engines were born... if they had had to ask each site for permission to read, rework and provide the "external" search service, everything would have died immediately (at least in Europe and use )
another example is allowing emerging countries to maintain lighter copyright laws to facilitate their growth (yes, they violate them, yes perhaps a small percentage of Zambian inhabitants would perhaps buy the media by paying Western fees, but it is better to leave it alone and look at a greater good than managing copyright guarantees everywhere...) the Americans interpret this legislative concept better, the Europeans if they don't wake up will be left behind for a long time,
and I repeat, comparing the passing of a dataset containing copyrighted material into matrices for the generation of an AI is very different from slavery
> As long as you can get over the synthetic data event horizon, where the model is smart enough to make good synthetic data, everything will be fine,” Mr. Altman said.
This seems intuitively incorrect to me. How can you advance the capabilities of a model by using the output of a model of the same level?
LLMs have no feedback loop with reality if it is short circuited with itself. It needs fresh training data.
english speaking rate: ~150 words per minute
gpt tokens per english word: ~1.3 tokens per word
1M hours = ~12B tokens
I'd fudge that up a bit because it's probably more tokens from non-English but comparatively seems like not much
The terms of service though? who cares? deactivate their google account then, they can still access the videos unauthenticated.
Personally I’ve made so much money off of AI generated works that would not be protected by copyright at all, that I think creators and owners of copyrighted works should re-evaluate their relationship with copyright. It seems more like copyright gives a sense of pride and accomplishment while almost no work is being monetized by its copyright or license alone.
but I would say worrying about protecting that right and “doing it right” is more work than actually having a license agreement or court filing, both of which are rare
It is liberating to be told “this can’t have copyright” and make enough money from entertaining the humans, who cares if it gets replicated and why didnt they just AI generate their own thing
GPT-3.5: ?
GPT-4: YouTube
GPT-5: ???
Perhaps capture of audio and video from places lacking legal expectation of privacy.
Buying conversations from listening devices at live nation sports venues or clear channel airports.
Spooky, but probable.
Maybe there are troves of data yet to be unearthed, or needing extraction from legacy media formats. Orgs aren't desperate enough to do that kind of schlep, but soon enough.
Like archives of terrestrial AM/FM or even HAM radio up for grabs. Going backward to go forward.