Also, it's possible that OpenAI is still training GPT-4, perhaps with additional modalities, and will make future snapshots available as public releases.
Also, it's possible that OpenAI is still training GPT-4, perhaps with additional modalities, and will make future snapshots available as public releases.
Maybe true, but he also said "We are not here to jerk ourselves off about parameter count"
https://techcrunch.com/2023/04/14/sam-altman-size-of-llms-wo...
Also, who says that the "transformer scaling laws" are the ultimate arbiter of LLM scaling? They overturned previous scaling laws and other scaling laws might overturn them. Furthermore, it's even possible that the transformer model won't even be used in later models. I remember Ilya making the point that just because the transformer model was the first one that looks like it can scale intelligence just by lighting up billions of dollars of GPUs, it doesn't mean it's the last one. Maybe it will even be like, the vacuum tube of AI models, and other ones are being made in secret. A hacker news rumor was that they are paying $5M-$20M per year to the top neural net experts probably to make some exotic architectures to surpass transformer.
This reminds me a TV interview of the author Patrick Modiano, just after he won the literature Nobel price. The presenter asked him if the money would help. The author answered essentially that the next time he would be in front of a white page, the money surely wouldn't help.
In the case of surpassing transformers, money could help to give access to more compute power. It could also help to prevent the research from being public.
As always, wealthy people and their "money doesn't make happiness" bullshit.
Can you link to a comparison or graph of obsolete and new scaling laws?
5 years at $20 million equals the highest paid FIFA player’s single year pay (or Tom Cruise’s pay for one movie, Top Gun: Maverick), or about two years of the highest paid NFL players, so it won't catch up the their wealth over that time (assuming similar lifestyle) as its losing ground to them every year.
(And the claim was total, not per person, salary, anyway.)
99% chance it's made up.
That said, if they thought a specific individual had even a reasonable chance of coming up with an improvement on the current state-of-the-art AI architecture that they'd be able to keep entirely to themselves, $20M would be a massive bargain.
The rumor is still almost certainly fake, but for someone very specific at this critical time in the field, I don't know if the number would be that absurd.
I’d pay creator of GPT that money easily. Probably not anyone else
I guess the reason why AI is so interesting is that human stupidity is so widespread.
There are also all of the quantization and other tricks out there.
Also they have demonstrated that the model already understands images but just haven't completed the API for this.
So they use quantization to increase the speed by a factor of 3 while slightly increasing the parameter count. Maybe find a way to make the network more sparse and efficient so in the end with the quantization the model actually uses significantly less memory. and continue with the RHLF focusing on even more difficult tasks and those that incorporate visual data.
Then instead of calling it GPT-5 they just call it GPT-4.5. Twice as fast as GPT-4, IQ goes from 130 to 155. And the API now allows images to be passed in and analyzed.
It's amazing the extreme levels of advantage that groups have depending on funding and connections.
It's actually not.
For now I'm using models like Salesforce/blip2 and OVF and Meta's Segment Anything for visual questioning.
It’s a very limited model for a select few.
hmm this is new. source for the image generation piece?
ChatGPT is IMO a heavily fine-tuned Curie sized model (same price via API + less cognitive capacity than even text davinci-003) so it would make sense that a heavily fine-tuned Davinci sized model would yield similar results to GPT-4.
I'd expect that by now we would enjoy similar speeds but this hasn't yet happened.
If they still haven't implemented these, it would be positively surprising (to me) to see the model run at similar speeds as chatgpt now. It'd be a great achievement if they really packed such performance on similar architecture (say by just training longer)
If you have chatGPT Plus you can choose "Legacy" from the drop-down to get the smarter (and slower) 175B Parameter version of GPT-3.5. That version is the same speed as GPT-4 when load is low (early morning EST), which lends credence to the theory that GPT-4 is the same size as overparametrized GPT-3.
It’s been established that LLMs are sensitive to corpus selection which is part of why we see anecdotal variance in quality across different LLM releases.
While we could increase the corpus of text by loading social media comments, self published books, and other similar text - this may negatively impact final model quality/utility.
Personal opinion, not OAI/GH/MSFT’s
Read OpenAI API docs on GPT model versions carefully, and look at them again from time to time.