Of course not to imply that OpenAI is the same as FTX.
It was just a rank and file IC.
At this point Gemini is as good as GPT4 (and anecdotally I think it's better at many coding assistance tasks, based on my experiments on the LMSys chatbot arena). Sora is getting a lot of press for text-to-video, but Google's Lumiere has already been out for a few weeks and produces pretty good results.
I have no doubt that OpenAI has been cooking things up. GPT-4 is an older model which others are only just catching up to now. They have an A+ team, a giant war chest, and a lot of momentum. But just because they have momentum does not mean that they have a moat.
I'd be willing to wager that the top-of-the-line foundational models are going to converge and become indistinguishable for almost all tasks. Even the open foundation models (e.g. Mixtral) are getting really good. Foundation models are not moats.
The players that have moats are Microsoft (with deep enterprise software & B2B expertise, allowing them to sell AI-powered software & ward off competition from upstarts) and Google (where their decade of investment into custom silicon allows them to train & run inference for cheaper than anyone else, by far).
10m context with that retrieval rate is such a monstrous leap. And to top it off, we got LargeWorldModel in the same week, capable of 1M token context with insane retrieval rate in the open source space. So not only is the open source world currently technically ahead of ChatGPT, so is Google. Which is why they had to announce SORA, because google's model is so far ahead of the competition. That's also why it will probably be ages before we get access to SORA. Now don't get me wrong, the average person can't afford 32 TPU's to run LWM, but we already have quants for it, which is a step towards enabling the average person (that somehow has 24-48gb of VRAM to get a taste of that power).
What is also striking is the fact that the new models are all multimodal as a standard. We not only leapfrogged in context size, but also in modalities. The model seems to only benefit from having more modalities to work with.
I think the statement Bill Gates made claiming that "LLM's have reached a plateau" itself indicates they don't believe they can make more money from training better/larger models. Which indicates that they already did as well as they could with their existing people, and are now "years" behind google. I never thought google could catch up, especially after their infamous "We have no moat" situation. But it seems they actually doubled down and did something about it.
To a lot of people, last Thursday was a very nihilistic day for Local Models, as the goalposts shifted from 128-200k context to 10M tokens with near perfect retrieval. It's literally insanely scary. But luckily we got LWM, and that means we have only been 10xed.
Now the local people will work on figuring out how to bridge the gap, before being leapfrogged again. What is really insane is that, we have had LLAMA2 for over a year now, and nobody else figured out how to get this result from it, despite it being around so long.
I still believe there are modifications to the architecture of MoE that will unlock new powers that we haven't even dreamed of yet.
Sorry, this was supposed to be well thought out, but it turned more into stream of consciousness, and I honestly had no intention of disagreeing with you.
If I remember the paper correctly, it was something about a 4M context in there. So not 10x, but 2.5x.
> What is really insane is that, we have had LLAMA2 for over a year now, and nobody else figured out how to get this result from it, despite it being around so long.
This isn't true. For now, the task of extending context to 10M tokens is brute-forced by money (increased HW requirements for training and inference and increased training time are also a financial domain). And for now, there simply is no leapfrogging solution for open source or commercial models, which will decrease the costs by orders of magnitude.
I literally have a 24/7 consultant with surface level understand of any topic in human history. This consultant also happens to be an amazing artist for $25 a month that gives me the rights to commercialize any art piece they make for me.
This is extremely hard to beat. This will be extremely hard to beat.
As a business they are succeeding.
LLMs are useful tools, but if they don't lead to something you can actually somewhat rely on to generate correct/consistent/factual results (see e.g. the recent Air Canada chatbot lawsuit) then the hype is a bust. TBD.
The bar I'm setting is far from "impossible"; even human children generally won't seamlessly confabulate when you ask them a question they don't know the answer to. Again citing the recent Air Canada case, these models can't even reliably answer simple questions that are definitively and objectively answered in documentation that is presumably made as freely available to them as is technically possible under the limitations of current technology.
No well-informed sources are saying that. If you're saying that mainstream reporting and other non-tech folks are wrong about what the possibilities are then... obviously.
Yes they are, actually! I've talked to people who obsessively read practically every LLM paper that passes through the arxiv, with undeniably deep and broad knowledge of the current state of this tech, who seriously believe it's going to surpass humans within a year or two. That it may already have, in the deep dark top secret labs beneath OpenAI HQ.
However,
> If you're saying that mainstream reporting and other non-tech folks are wrong about what the possibilities are then... obviously.
If it was obvious then why are you still replying to my comments, which have very obviously been specifically addressing the mismatch between hype and reality? If the current approach doesn't scale, the hype will have been a bust! Objectively! That's what I've been talking about this whole time!
"Mainstream reporting and other non-tech folks" is a bit disingenuous, though. The primary drivers of the current unrealistic hype are software vendors and associated clingers-on looking to make a quick buck. They'll say anything, regardless of whether it's true, and as a result of those mostly-falsehoods our public lives will be flooded with awful AI tools that make everything shittier and more difficult. I can't wait!
Like, are you following the things that OpenAI and Deepmind are saying at all? The things that make current LLMs not a threat, they aim to tear down as soon as they can arrange.
OpenAI just released a video network, and one of their core touted benefits was that you could use it as an action controller!
And, um. Do you really think, when the AIs can take a simple prompt and turn it into a ten thousand step plan that requires dynamic skill acquisition, resourcing and persistence, that generating the prompt will be the one single task that stumps them? When we are at that point - and to be clear, every leading AI organisation is sprinting to reach that point earliest - then the difference between doom and safety will be one sentence: "When you are done with that, generate a new prompt." This is not how a world with a long expected lifespan looks.
edit: To be clear, I'm still not accusing OpenAI of making up the doom stuff. Even though when I phrase it like that it sounds like they're directly working on things that obviously end the world, which seems contradictory, I don't think they see it like that. To be honest, I can't explain why any doomer works at OpenAI, except in the way that people sometimes move towards explosions and gunfire. I think it's just a bug in the human brain. We want to have the danger in sight.
So I just don't think this captures an important distinction at the limit. If a system can generate a good action plan, turning it into an agent is just plumbing.