Sam Altman: OpenAI is not training GPT-5 and "won't for some time"
twitter.com
twitter.com
Also, it's possible that OpenAI is still training GPT-4, perhaps with additional modalities, and will make future snapshots available as public releases.
Personal opinion, not OAI/GH/MSFT’s
ChatGPT is IMO a heavily fine-tuned Curie sized model (same price via API + less cognitive capacity than even text davinci-003) so it would make sense that a heavily fine-tuned Davinci sized model would yield similar results to GPT-4.
I'd expect that by now we would enjoy similar speeds but this hasn't yet happened.
If they still haven't implemented these, it would be positively surprising (to me) to see the model run at similar speeds as chatgpt now. It'd be a great achievement if they really packed such performance on similar architecture (say by just training longer)
If you have chatGPT Plus you can choose "Legacy" from the drop-down to get the smarter (and slower) 175B Parameter version of GPT-3.5. That version is the same speed as GPT-4 when load is low (early morning EST), which lends credence to the theory that GPT-4 is the same size as overparametrized GPT-3.
There are also all of the quantization and other tricks out there.
Also they have demonstrated that the model already understands images but just haven't completed the API for this.
So they use quantization to increase the speed by a factor of 3 while slightly increasing the parameter count. Maybe find a way to make the network more sparse and efficient so in the end with the quantization the model actually uses significantly less memory. and continue with the RHLF focusing on even more difficult tasks and those that incorporate visual data.
Then instead of calling it GPT-5 they just call it GPT-4.5. Twice as fast as GPT-4, IQ goes from 130 to 155. And the API now allows images to be passed in and analyzed.
It's amazing the extreme levels of advantage that groups have depending on funding and connections.
It's actually not.
For now I'm using models like Salesforce/blip2 and OVF and Meta's Segment Anything for visual questioning.
It’s a very limited model for a select few.
hmm this is new. source for the image generation piece?
Maybe true, but he also said "We are not here to jerk ourselves off about parameter count"
https://techcrunch.com/2023/04/14/sam-altman-size-of-llms-wo...
Also, who says that the "transformer scaling laws" are the ultimate arbiter of LLM scaling? They overturned previous scaling laws and other scaling laws might overturn them. Furthermore, it's even possible that the transformer model won't even be used in later models. I remember Ilya making the point that just because the transformer model was the first one that looks like it can scale intelligence just by lighting up billions of dollars of GPUs, it doesn't mean it's the last one. Maybe it will even be like, the vacuum tube of AI models, and other ones are being made in secret. A hacker news rumor was that they are paying $5M-$20M per year to the top neural net experts probably to make some exotic architectures to surpass transformer.
This reminds me a TV interview of the author Patrick Modiano, just after he won the literature Nobel price. The presenter asked him if the money would help. The author answered essentially that the next time he would be in front of a white page, the money surely wouldn't help.
In the case of surpassing transformers, money could help to give access to more compute power. It could also help to prevent the research from being public.
As always, wealthy people and their "money doesn't make happiness" bullshit.
Can you link to a comparison or graph of obsolete and new scaling laws?
5 years at $20 million equals the highest paid FIFA player’s single year pay (or Tom Cruise’s pay for one movie, Top Gun: Maverick), or about two years of the highest paid NFL players, so it won't catch up the their wealth over that time (assuming similar lifestyle) as its losing ground to them every year.
(And the claim was total, not per person, salary, anyway.)
99% chance it's made up.
That said, if they thought a specific individual had even a reasonable chance of coming up with an improvement on the current state-of-the-art AI architecture that they'd be able to keep entirely to themselves, $20M would be a massive bargain.
The rumor is still almost certainly fake, but for someone very specific at this critical time in the field, I don't know if the number would be that absurd.
I’d pay creator of GPT that money easily. Probably not anyone else
I guess the reason why AI is so interesting is that human stupidity is so widespread.
Read OpenAI API docs on GPT model versions carefully, and look at them again from time to time.
It’s been established that LLMs are sensitive to corpus selection which is part of why we see anecdotal variance in quality across different LLM releases.
While we could increase the corpus of text by loading social media comments, self published books, and other similar text - this may negatively impact final model quality/utility.
Also, the GPT-4 message cap at chat.openai.com was shown as something along the lines of "we expect lower caps next week", then changed to "expect lower caps as we adjust for demand" to "GPT-4 currently has a cap of …". This sounds to me like they changed from having lots of compute to being limited by it. Also note how everything at OpenAI is now behind a sign up and their marketing has slowed down. Similarly, Midjourney has stopped offering their free plan due to lack of compute.
Seems like we didn’t need a 6 months pause letter. Hardware constraints limit the progress for now.
So, I think there's a lot of improvements where maybe gpt-4, is as far you go in terms of inputting data, and maybe better use cases are more customization of data trained on, or finding ways of going smaller, or even some model that just trains itself on the data requirements, similar to how we jump on google when we're stuck, it'd do the same and build up its knowledge that way.
I also think we need improvements in vector stores that maybe add weights to "memories" based on time/frequency/recency/popularity.
At the time I noticed that the wording they gave technically implied they expected the cap to get more limiting and then that's exactly what happened, and I haven't been able to work out if that is indeed what was the intended message or not.
I asked it to help me code something. Then it stopped midway through, so I asked it to continue from the last line.
…It started from the beginning.
Now at the same point, I asked it not to stop. To keep going.
It started again from the beginning.
It went like this for about another 10 or so prompts. Hell, I even asked it to help me write a better prompt to ask it to continue from the line it cut off and I then used that. It didn’t work at all.
Then I ran out of prompts.
Three hours later, it did the same crap to me and I lost around 14 prompts to it being ‘stuck’ in an eternal loop.
Basically, OpenAI are sneaky devils. ‘Stuck’ my ass - that was intentional to free up resources.
Where did you read that? You’d think it would be there for people to read right next to the prompt.
The game theory behind AGI research is identical to that of nuclear weapons development. There exists a development gap (the size of which is unknowable ahead of time) where an actor that achieves AGI first, and plays their cards right, can permanently suppress all other AGI research.
Even if one's intentions are completely good, failure to be first could result in never being able to reach the finish line. It's absolutely in OpenAI's interest to conceal critical information, and mislead competing actors into thinking they don't have to move as quickly as they can.
Nuclear powers have not been able to reliably suppress others from creating nuclear weapons. Why would we think the first AGI will suppress all others perfectly?
What is clear is that on this earth itself we have cetacean, corvid, cephalopod intelligence which is wired very differently. Perhaps we need to respect the diversity of intelligences that exist and study this growth in LLM and adjoint areas as just synthetic intelligence.
Rebranding maybe could help drive a level of objectivity this conversation on ethics etc that seems to be missing
It could be useful for a similar reason as the euphemism treadmill. We could leave behind all of the misguided assumptions about AI with the old 'artificial intelligence' nomenclature and move forward with 'synthetic intelligence' which has our new understanding of what systems like GPT-4 can do.
Since nobody actually knows what "intelligence" is, the word will mean to people whatever they want it to mean.
Everybody knows what intelligence is. Even if we can't agree on a precise definition, it's pretty obvious that it's the thing that humans and other animals do that involves learning, reasoning, planning, and problem solving. We can also agree that being successful at certain tasks constitutes intelligence. Solving a math problem is intelligence. Writing a poem is intelligence.
Much like...
"Everyone knows what porn is"
"Everyone knows who god is"
"Everyone knows what beauty is"
The devil is in the details and rather generic words that describe a gradient can never capture the exact nature of what we're trying to define in specific situations.
This sort of problem is common with language, and is a great example of why I'm not really on board with using natural language for technical things.
I pretty much agree with that, so I'm not sure where the disagreement is here. Let me go back to the original statement I was responding to.
>Since nobody actually knows what "intelligence" is, the word will mean to people whatever they want it to mean.
If I tell you someone is intelligent, you roughly know what I am talking about. Just because it's hard to formalize that doesn't mean that that the word can mean whatever people want it to mean. For example, if I tell you my friend is intelligent, you would be wrong to interpret that as meaning that my friend has red hair, because hair color is irrelevant to the traits that we normally associate with intelligence. The fact that there are right and wrong ways of interpreting my sentence implies that there is some generally agreed upon notion of what intelligence is, even if that notion is fuzzy and has grey areas.
I'm not sure we are disagreeing. I'm just having a discussion.
> If I tell you someone is intelligent, you roughly know what I am talking about.
Correct, because the context (you're talking about a human, and I know roughly what that means with humans) narrows the possibilities. But even there, it's a vague sort of intuitive knowledge, like trying to say what "art" is.
But when it comes to other areas -- such as machines -- context doesn't help narrow the possible meanings. What does saying a machine is "intelligent" mean? If you ask a machine learning person, you'll get a reasonably specific answer. If you ask the average person on the street, you'll get very, very different answers.
The reason is because we don't know what "intelligence" actually is. We don't even know, with any specificity, what it is in humans -- which is why psychologists assert that there are multiple kinds of intelligence (even if they disagree about how many there are).
> even if that notion is fuzzy and has grey areas.
I don't disagree at all. But the notion has more fuzzy and gray areas than solid ones. As an example, when most people imagine an "artificial intelligence", what they're really imagining is "consciousness". Is consciousness required for intelligence? Who knows? The answer to that depends on what you mean by "intelligence" and we don't agree enough on what that means to have that sort of discussion without beginning by defining the terms.
I don't disagree with this. What I'm saying is that that definition, while reasonable and I agree, is one that we've just decided on for this conversation.
It isn't one that would be considered complete and correct in all discussions about intelligence.
> I don't think the implementation details matter.
I agree, for the definition of intelligence you just cited. But my point is that "intelligence" is not well-defined or understood. I'm genuinely surprised that people think this is a controversial stance -- I really thought it was well-understood.
We can settle on a definition for rhetorical purposes (and, I would argue, that's mandatory in order to have any solid discussion about intelligence), but any definition we agree on will leave out a lot of things that people consider part of "intelligence".
Only if you care about those details. Almost no one does.
In almost any conversation, everyone does in fact know what intelligence, porn, god and beauty are. Yes, all those ideas are fuzzy at the borders, but we almost never need to resolve them in detail when talking about them. When we do, then yes, things get tricky and there's a lot of disagreement - but at the end of the day, as the phrase I once read on the Internet goes, it all has to add up to normality. You can still work with fuzzy, casual concepts, even though you can't define them precisely.
> We can also agree that being successful at certain tasks constitutes intelligence. Solving a math problem is intelligence. Writing a poem is intelligence.
As an example, I don't agree that either of those things indicates intelligence all by themselves. We've had programs that nobody would call "intelligent" to do both of those things for decades.
So you're right that if I have separate algorithms, each designed for a specific purpose, that those algorithms aren't intelligent. However, if I have a general system that can learn how to solve a math problem, write a poem, and do a bunch of other things that humans can do, then that system is intelligent.
AI is really an overloaded term that includes 70 years of snake oil, Skynet, the Singularity and killer robots. I think we need a new name to start fresh.
And personally, I think we are extremely biased by our sci-fi to think of this tech as malevolent. As far as we can see, it can only know what we teach it since it relies on all of our perceptions to learn. LLMs seem both extremely promising as a useful tool and very pliant to the operator’s wishes. I’m way beyond “this is a fancy next word predictor” as I think it’s emergent behavior has many of the hallmarks of reasoning and novel inference, but at best I think it is only part of a mind and an unconscious one at that.
https://blog.quintarelli.it/2019/11/lets-forget-the-term-ai-...
> I think it’s emergent behavior has many of the hallmarks of reasoning and novel inference
You contradicted yourself here, right?
I quite like the term, and it seems quite unique (perhaps cribbing from 'grounded cognition' though that's an entirely different idea AFAIK)
Do you have a word that you prefer?
> Do you have a word that you prefer?
Nope. That's why I use "intelligence" despite the problems with it. "Intelligence" may be a blank slate on which you can write whatever meaning you wish, but at least it can stretch to mean something accurate.
'Understanding' has an ambiguous meaning, while thought and senses are certainly not applicable.
But it's a semantic discussion about a novel and not well understood topic so :shrug: :)
I agree entirely.
[1] https://www.theverge.com/2023/1/23/23567448/microsoft-openai...
Microsoft invested $10B, and get 75% of OpenAI’s profits until they receive $10B and _then_ they get their 49% stock certificate.
[0] https://www.semafor.com/article/01/09/2023/microsoft-eyes-10...
But if I do this in a stock context and buy 49% control of multiple companies over and over, with all the same obviousness of my intentional avoidance of the regulatory trigger, it's considered a smart move and pretty much the status quo.
Yes, the practice of law says Microsoft does not own openai. But it's also obvious what's going on when companies do this.
You can go down the rabbit hole if you want, but if you want only the most superficial glimpse of it then consider that OpenAI board member Will Hurd was a CIA undercover agent and also a representative in the House Permanent Select Committee on Intelligence and also he is a trustee of In-Q-Tel which is the private investment arm of the CIA.
Buddy, these guys are so far behind the times, they're constantly playing catch-up from 10-20 years ago.
Eric Schmidt, former Google CEO, led a multi-year project to develop a national AI strategy, https://www.nscai.gov/
> The Final Report presents the NSCAI’s strategy for winning the artificial intelligence era. The 16 chapters explain the steps the United States must take to responsibly use AI for national security and defense, defend against AI threats, and promote AI innovation. The accompanying Blueprints for Action provide detailed plans for the U.S. Government to implement the recommendations.
That's pretty much exactly how one of the OpenAI Red Teamers Nathan Labenz describes the raw GPT-4, starting around 45 minutes into the video:
That's not really how he described it.
His point is that the raw model that became GPT-4 would do literally anything it asked you to.
It would write fascist propaganda just as readily as it would offer medical advice. Literally any and all input from the user was fair game.
But it wouldn't just veer from medical advice into fascist propaganda, not unless the user was steering it in that direction.
It's unfortunate that people would abuse that and we can't have the raw model just for personal use. The story telling and characters alone would be worth it. The safety guards tend to seep into fictional scenarios, making them more bland and preachy.
When manifest as ChatGPT, it is obvious that what presents as 1 magical solution is in fact an elaborate combination of varying degrees of innovation.
In my view, the reasoning for not releasing GPT4 information (hyperparameters, etc) had nothing to do with AI safety. It was a deliberate marketing decision to obscure how the sausage is actually made.
In their technical report they give both reasons:
"Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."
On the one hand they won’t talk details because it’s about safety, then open ai just throws the thing on the Internet, with “tools”
Then it’s about the competitive landscape, which is completely ridiculous as you can’t be competing to building more powerful AIs in secret by means of an arms race and respect safety protocols and be “careful”. Realistically, the top minds , and thinkers of our time should be working carefully and together to develop tech like this. Not competing against each other in secret labs.
There’s just way too many mixed messages in their story. They seem confused themselves about what they’re trying to achieve. Which I find unsettling. Now I can’t help but think like everything else, this is just all banking hoping to establish a monopoly via regulation and raise as much money as possible.
The connectivists left the 'AI' folks and established the ML field in the 90s.
Sometimes those political rifts arise in discussions about what is possible.
Thinking of ML under the PAC learning lens will show you why AGI isn't possible through just ML
But the Symbolists direction is also blocked by fundamental limits of math and CS with Gödel's work being one example.
LLMs are AI id your definition is closer to the general understanding of the word, but you have to agree on a definition to reach agreement between two parties.
The belief that AGI is close is speculative and there are many problems, some which are firmly thought to be unsolvable with current computers.
AGI is pseudo-science today without massive advances. But unfortunately as there isn't a consensus on what intelligence is those discussions are difficult also.
Overloaded terms make it very difficult to have discussions on what is possible.
'GPT-4 is not “AI” because AI means “AGI,”'
Is a more strict term that hasn't typically applied to AI as an example of my above claim.
As we lack general definitions, it isn't invalid but no AI is thought to be possible with their claims.
AI being computer systems that perform work that typically requires humans within a restricted domain is closer to what most researchers would use in my experience.
An AI's ability to contemplate life while sitting under a tree is secondary to the impact it has on society.
one pattern in modern history has been that communication and creative technologies, be it the television or even the internet, had significantly less economic impact than people expected. Both television and the internet may have transformed business models and had huge cultural impact, but all things considered negligible impact on total productivity or the physical economy.
Given that generative AI seems much more suited to purely virtual tasks or content creation than physical work I expect that to repeat. Over past cycles people have vastly overstated rather than underestimated the impact tech has on the labor force.
It is obscenely hard to measure those sorts of things.
OK, here is one example.
It used to be that engineers (the physical kind, not the software kind) had their own secretaries to manage meetings and fetch documents.
Those secretaries are all out of work now, replaced by Outlook and PDFs.
Modern farms are wired with thousands upon thousands of IOT sensors, precisely controlling every aspect of the fields and crops. Soil is maintained in perfectly ideal conditions. The internet is what made this possible.
The Internet has allowed for individuals to easily trade stocks, which has had who knows how large of an impact on the economy, but I am willing to guess it isn't a small one.
The Internet also enabled all sorts of algorithmic trading to pop up.
Television is of course a huge source of economic output in its own right.
6.9% of the US GDP is Media and Entertainment, not sure if that includes video games or not.
The tech industry is at least 10% of the US GDP, remove the Internet and that drops dramatically.
Are you only talking about consumer products like YouTube or are you including everything related to the unparalleled exchange of data globally?
I cannot imagine the impact of a global information network for businesses having less impact than a few x2’s on just about every relevant axis you can imagine.
Why? PAC looks a lot like how humans think
> But the Symbolists direction is also blocked by fundamental limits of math and CS with Gödel's work being one example.
Why? Gödel's incompleteness appoes equally well to humans as machines. It's an extremely technical statement about self-reference within an axiom systems, pointing out that it's possible to construct paradoxical sentences. That has nothing to do with general theorem proving about the world.
Gödel's incompleteness applies equally well to humans as machines, in writing down axioms and formula, not in general tasks.
The Irony of trying to explain this on a site called Y combinator, but even for just prepositional logic, exponential-time is the best that we can do for for algorithms and general proof tasks.
For first order predicate logic, finding valid formulas is recursively enumerable, thus with unlimited resources they can be found in finite time.
But unlimited resources and finite time are not practical.
Similar with modern SOTA LLMs, while they could be computationally complete they would require unbounded amount of ram to do so, which is also impractical. Also invalid formula cannot reliably be detected.
Why this is ironic.
The Curry's Y combinator: Y = λf.(λx.(x x)) (λx.(x x)), lead to several paradoxes show that untyped lambda calculus is unsound as a deductive system.
Church–Turing thesis shows that lambda calculus and Turing machines are equivalent.
Here is Haskell Curry's paper on the Kleene–Rosser paradox which is related.
https://www.ams.org/journals/tran/1941-050-03/S0002-9947-194...
The way I know the story is that modern machine learning started as an effort to overcome the "knowledge acquisition bottleneck" in expert systems, in the '80s. The "knowledge acquisition bottleneck" was simply the fact that it is very difficult to encode the knowledge of experts in a set of production rules for an expert system's knowledge-base.
So people started looking for ways to acquire knowledge automatically. Since the use case was to automatically create a rule-base for an expert system, the models they built were symbolic models, at least at first. For example, if you read the machine learning literature from that era (again, we're at the late '80s and early '90s) you'll find it dominated by the work of Ryszard Michalski [1], which was all entirely symbolic as far as I can tell. Staple representations used in machine learning models of the era included decision lists, and decision trees, and that's where decision tree learners, like ID4, C45, Random Forests, Gradient Boosted Trees, and so on, come; which btw are all symbolic models (they are and-or trees, propositional logic formulae).
A standard textbook from that era of machine learning is Tom Mitchell's "Machine Learning" [2] where you can find entire chapters about rule learning, decision tree learning, and other symbolic machine learning subjects, as well as one on neural network learning.
I don't think connectionists ever left, as you say, the "AI" folks. I don't know the history of connectionism as well as that of symbolic machine learning (which I've studied) but from what I understand, connectionist approaches found early application in the field of Pattern Recognition, where the subject of study was primarily machine vision.
In any case, the idea that the connectionists and the symbolists are diametrically opposed camps within AI reserach is a bit of a myth. Many of the luminaries of AI would have found it odd, for example Claude Shannon [3] invented both logic gates and information theory, whereas the original artificial neuron, the Pitts and McCulloch neuron, was a propositional logic circuit that learned its own boolean function. And you wouldn't believe it but Jurgen Schmidhuber's doctoral thesis was a genetic algorithm implemented in ... Prolog [4].
It seems that in recent years people have found it easier to argue that symbolic and connectionist approaches are antithetical and somehow inimical to each other, but I think that's more of an excuse to not have to learn at least a bit about both; which is hard work, no doubt.
______________
[1] https://en.wikipedia.org/wiki/Ryszard_S._Michalski
[2] It's available as a free download from Tom Mitchell's wesbite:
http://www.cs.cmu.edu/afs/cs.cmu.edu/user/mitchell/ftp/mlboo...
[3] Shannon was one of the organisers of the Dartmouth Convention where the term "Artificial Intelligence" was coined, alongside John McCarthy and Marvin Minsky.
[4] https://people.idsia.ch/~juergen/genetic-programming-1987.ht...
The second question in particular drops the potential implicit assumption that only 100 people stand in line each day.
I face this issue in my CS masters program constantly, and would probably have failed this test much the same as GPT did.
Turing's test was not "this computer fooled me over text, therefore it's an AI". It's a philosophical, "we want to consider a machine that thinks, well we can't really define what thinking is, so instead it's more important to observe if a machine is indistinguishable from a thinker." He then goes on to consider counterpoints to the question, "Can a machine think?" Which is funny because some of these counterpoints are similar to the ones in the author's article.
Author offers no definition of "think" or "invent" or other words. It's paragraph after paragraph of claiming cognitive superiority. Turing's test isn't broken, it's just a baseline for discussion. And comparing it to SHA-1 is foolish. Author would have done better with a writeup of the Chinese room argument.
in before comments on auto gpt.
Realistically though, the road to AGI and beyond is like the expansion of the human race to The Moon, Mars and beyond, slow, laborious, capital and resource intensive with vast amounts of discoveries that still need to be made.
Edit: Off the top of my foggy head, LLMs as I understand them are text completion predictors based on statistical probabilities trained on vast amounts of examples of what humans have previously written, whose output is styled with neuro linguistic programming also based on vast numbers of styles of human writing. This is my causual amatuer understanding. There is no logical, reasoning programing such as the Lisp programmers attempted in the 1980's, but clearly the logical abilities of the current LLMs fall short and they are not AGI for that reason. So how do we add logic abilities to make LLMs AGI? Should we revisit the approaches of the Lisp machines of the 1980's? This requires much research and discovery. Then there's the question of just what is general intelligence. I've always thought that emotional intelligence played a huge role in high intelligence, a balance between logic and emotion or Wise Mind is wisdom. Obviously we won't be building emotions into silicon machines or will we? Is anyone proposing this? This could take hundreds of years to accomplish if it is even possible. We could simulate emotion but that's not the same, that's logic. Logical intelligence and emotional capability I think are a prerequisite for consciousness and spirituality. If the Universe is conscious and it arises in a focused manner in brains that are capable of it then how do we build a machine capable of having consciousness arise in it? That's all I'm saying.
The human brain uses on the order of 10 watts of power and there are almost 8 billion examples of this. So we have hard proof that from a thermodynamic perspective general intelligence is utterly and completely mundane.
We almost certainly already have the computational power required for AGI, but have no idea what a complete working architecture looks like. Figuring that out might take decades, or we might get there significantly quicker. The timespan is simply not knowable ahead of time.
I'm not concerned in the slightest about "the singularity" and non-aligned superintelligences. AGI in the hands of malicious human actors is already a nightmare scenario.
Tracking problematic actions back to the person that own the AGI will likely not be a difficult task. The owner of an AGI would be held responsible for its actions. The worry is that these actions would happen very quickly. This too can be managed by safety systems, although they may need to be developed more fully in the near future.
And actually there is no law or anything that says that any particular change or improvement to the model or even new training run that necessities them calling it version 5. It's not like there is a Version Release Police that evaluates all of the version numbers and puts people in jail if they don't adhere to some specific consistent scheme.
source?
> "it’s easy to create a continuum of incrementally-better AIs (such as by deploying subsequent checkpoints of a given training run), which presents a safety opportunity very unlike our historical approach of infrequent major model upgrades."
Of *course* they’re still training chat-gpt5
It’s just grokking an image directly. Were the pixels tokenized somehow? I’m very curious what that does to a model like this.
Can somebody that actually knows anything clue me in?
Multi-model text-image transformers add these tokens right beside the text tokens. So there is both transfer-learning and similarity graphed between text and image tokens. As far as the model knows they're all just tokens. It can't tell the difference between the two.
For the model, the tokens for the words blue/azure/teal and all the tokens for image patches with blue are just tokens with a lot of similarity. It doesn't know if the token its being fed is text, image, or even audio or other sensory data. All tokens are just a number with associated weights to a transformer, regardless of what they represent to us.
The GPT-4 vision API is actually in production and in at least two public products already. https://www.bemyeyes.com/ and https://www.microsoft.com/en-us/ai/seeing-ai
blip-2(https://github.com/salesforce/LAVIS/tree/main/projects/blip2)
fromage(https://github.com/kohjingyu/fromage)
prismer(https://github.com/NVlabs/prismer)
palm-e(https://ai.googleblog.com/2023/03/palm-e-embodied-multimodal...)
now assuming gpt-4 vision isn't just some variant of mm-react(ie what you're describing), that's what's happening here. https://github.com/microsoft/MM-REACT
images can be tokenized. so what happens usually is that extra parameters are added to a frozen model and those parameters are trained on an image embedding to text embedding task. the details vary of course but that's a fairly general overview of what happens.
the image to text task the models get trained to do has its issues. it's lossy and not very robust. gpt-4 on the other hand looked incredibly robust. they may not be doing that. idk
Strictly speaking the model doesn't have to be frozen (though unfreezing tends to make the original model perform much worse at NLP tasks) and the task isn't necessarily just image to text (Palm e for example trains to extract semantic information from objects in an image as well)
I find most of the community models for Stable Diffusion to be pretty poor. A lot of models made from merges, waifu only models, models made from very small training datasets, and still stuck on 1.5 while 2.1 is much better overall.
However you have the amazing LoRas for 2.1 that are very powerful. The community is a bit ignoring them, but I think the potential is great.
But MidJourney is a lot better at text to image and image to image for now.
"OpenAI’s CEO says the age of giant AI models is already over": https://news.ycombinator.com/item?id=35603756 (shared 3 dags after this post)
Let's see what the future holds.
So maybe they’re working on a “system 2” now, which is perhaps more related to what deepmind is doing?
I wonder if there's an assumption for how big an LLM should be before it could even conceivably be an LLM. Is there a minimum size necessary before that capability is plausible?
As far as I know current LLMs are entirely static once trained, they don't learn at all in runtime.
From an engineering standpoint, even the less powerful GPT-3.5turbo model handles NLP tasks, really nice tools like LangChain and LlamaIndex that I covered in my last book make it easy to use your own data sources.
I think the possibilities of using what we currently have in useful projects are vast.
Regardless they can and will call future models anything they want. They could easily just decide that the minor improvements that come out in a few months are called GPT-4.2 and the major new training run is called GPT-4.5 instead of GPT-5.
So if some particular GPT-4 improved successor is based on the GPT-4 core transformer size and pretrained parameters then we'd call it GPT-4.x, but if some other GPT-4 successor is a larger core model (which inevitably also means it's re-trained from scratch) then we'd call it GPT-5, no matter if its observable performance is better or worse or comparable to the tweaked GPT-4.x options.
In terms of what this improvement would actually look like in terms of real world, emergent capabilities, no one knows.
“Will I still be able to feed my family, Sam? Ilya?”
The old capitalist treadmill is now on full speed and we’re all just trying to keep up.
Where it goes…nobody knows. What an entertaining drama.
What about Roko's basilisk?
They could train it with more data in the hopes of getting another big leap there, but what data is left? They've fed it everything it seems.
So what's left is getting the runtime reduced in terms of the model size. Hire some brilliant minds to turn an N-squared into an N-log-N (or something to that effect).
Maybe GPT4 has some ideas.
There is no 'revolution' around this. Just 'evolution' with more data and more excessive waste of compute to create another so-called AI black-box sophist with Sam Altman selling both the poison (GPT-4) and the antidote (Worldcoin).
At some point, with their tremendous lock-in strategy, O̶p̶e̶n̶AI.com and Microsoft will eventually use the lock-in to upsell and compete against their partners.
I actually find it pretty amazing that more people aren't given pause by Sam Altman's involvement here. After the WorldCoin stuff, I'd think that he'd be viewed with a much more skeptical eye in terms of his ethics.
Worldcoin also isn't an April fools joke either and they are quite serious in selling there proof-of-personhood anti-AI snake-oil antidote.
Why? It's clear that it positions Mr Altman to a point where he cannot lose and is hedged on both sides of where this AI narrative goes.
From my understanding they are training their new GPT models off of a checkpoint from the previous generation, so they technically have partially trained multiple future models in their GPT lineage.
The GPUs which are good for cryptocurrencies are decent for using ML models but are not good for training LLMs and vice versa, as the hardware requirements start to diverge. Training LLMs requires not only high compute power but also lots and lots of memory and extremely high-speed interconnect for large models, which ends up costing far more than the pure compute cryptomining needs, making it not cost-efficient for mining.