It can start posting synthesized ideas on social media and see how many likes it gets. Coupled with a metric containing dissimilarity to current information, this could be a useful way to progress to superhuman insights.
There was a rumor that they were going to use Whisper to transcribe YouTube videos and use that for training. Since it's multimodal, incorporating video frames alongside the transcriptions could significantly enhance its performance.
But I am quite sure that if you start doing it at scale, google will notice.
You could be sneaky, but people in this business talk (since they know another good paying job is just around the corner) so It would likely come out.
Google is also owns a lot of it's own backbone, so It would be a lot easier for them to play network games.
And they could even try to be sneaky and try poisoning the data if it comes to that.
And since OpenAI konw that, since they probably have people that used to work at google at some point, they are unlikely to try.
Even less likely if Microsoft would know. MS is probably the only company that has even more layers than Oracle and they would not approve.
This also makes me doubt that NSA hasn't already cracked this problem. Or that China won't eventually beat current western models since it will likely have way more data collected from its citizenry.
Then again it would give you data on every accent in the country, so the holy grail for modelling human speech.
The data is not finite.
That model might be very well tuned to solve IBM's internal problems.
I also question that most companies have the volume and quality of data worth training on. It's littered with cancelled projects, old products, and otherwise obsolete data. That's going to make your LLM hallucinate/give wrong answers. Especially for regulated and otherwise legally encumbered industries. Like can you deploy a chat bot that's wrong 1% or 0.1% of the time?
You have to understand that all the incentives are perfectly aligned for corporations to put this to work, even spending tens of millions in getting it right.
The first corporate CEO who announces that his company used AI to reduce employee costs while increasing profits is going to get such a fat bonus that everyone will follow along.
Maybe if you trained it on movies before CGI existed ?
> YouTubers upload about 720,000 hours of fresh video content per day. Over 500 hours of video were uploaded to YouTube per minute in 2020, which equals 30,000 new video uploads per hour. Between 2014 and 2020, the number of video hours uploaded grew by about 40%.
If a scrape of the general internet, scientific papers and books isn’t enough, a trillion trillion trillion text messages to mom aren’t going to change matters.
They might have trained on a lot of the 'high quality' tokens, however.
There's a ton of potential left on the table. The question is if transformers have hit their limit with GPT-4 or not.
It's a pretty simple equation when you think about it this way and why Sam would say they have hit their limit. Sam is basically Microsoft and they want to retain their lead. Once Google learns to put their data to use correctly, it's almost guaranteed game over for OpenAI if they want it to be.
Dataset size is not relevant to predicting the loss threshold of LLMs. You can keep pushing loss down by using the same sized dataset, but increasingly larger models.
Or augment the dataset using RLHF, which provides an "infinite" dataset to train LLMs on. Limited by the capabilities of the scoring model which, of course, you can scale the scoring model infinitely so again the limit isn't dataset size but training compute.
Deepmind and others would disagree with you! No-one really knows in actual fact.
[1] https://www.deepmind.com/publications/an-empirical-analysis-...
Merely pointing out that the debate as to whether we are compute or data limited (OP) has not concluded at all; There are lots of compelling theories on relationship between the two.