GPT-4’s secret
thealgorithmicbridge.substack.com
thealgorithmicbridge.substack.com
Do you have any real information about the architecture? Nope.
Do you have anything other than hearsay? Nope.
Do you know there isn’t actually some important breakthrough? Nope.
Is gpt4 still the best model available despite whatever shortcuts they may have taken? Yup.
…so… until we have some concrete, repeatable information, this is just a lot of hot air.
They did something.
No one knows what.
The results are better than what anyone else has done.
A bunch of people have speculations about what, but no one has been able to pull off an equivalent yet.
Same as yesterday.
whoa. archive.fo skips the substack paywall?
1. Firefox settings -> type dns in search bar
2. Under DNS over HTTPS either turn it off or alternatively select NextDNS as provider in the increased protection box.
3. archive today should start working now
welp...
If this is true, I wonder about how likely chain of thought is involved in passing data between the different models?
The 16 models in the ensemble are all trained in the same way, but attend to the input data differently. It’s a little like how multi-headed attention works by splitting the input embeddings into typically 8 parts and having a set of query, key, and value vectors each train on their own 1/8th slice of the input embeddings. Although there is no “meaning” to these slices, nonetheless the KQV vectors for each head will learn different relationships between the inputs simply because they were trained differently.
1. expanded from an outline (2-stage process)
or
2. being chained between models
If it was multi-stage, I'd expect some chinese whispers-style drift, where the plot gets lost slightly between steps. The responses I'm getting from GPT4 are focused and specific.
Likewise, if the responses were getting chained between models, I'd expect visible seams in tone / content in between.
My guess for how they're architecturing it, if it is 8 models, is either:
1. Each response handled by 1 model
or
2. Some kind of voting / confidence system that switches between models on the fly
GPT4 is definitely useful, but it points towards bad news for OpenAI and potentially the entire field. A lot of people really wanted to believe that OpenAI had some secret sauce that actually pointed towards a path at true AI.
Turns out they just poached Google's own researchers and did a better job at turning Google research into a product(all the papers for this type of architecture came from Google Brain, authors are now at OpenAI). OpenAI is doing impressive work on the practical side of AI but apparently nothing revolutionary in terms of research, which is why people are disappointed by this reveal.
What do you understand this to mean?
Basically it's a trick that recognizes that language transformers only perform computation to generate words so for complex tasks you can get better results by asking the model to explain its chain of thought and only give the answer at the end. This has the effect of giving the model "time to think". If it didn't generate those words it wouldn't have anything to hang that computation off since it is fundamentally a word-generation model.
Here's a simple example: keyterms from text are extracted with text detection from an image. Those keyterms will sometimes have bad reads where "aligned AI" might pop out as "aligned Al". A subsequent "internal thought" would be formed and ask, "What's wrong with the 'aligned Al' keyterm?". If an updated response is returned, we use it instead of the original output.
I think this is the paper that really kicked off this technique: https://arxiv.org/abs/2201.11903
You can google for more results / papers.
I think a more interesting article would be about sketching out technically some potential ways it could work in detail for an open source effort to try to imitate it. That could be a broad range of things because we don't even know at what level are the multiple things working together implemented.
Side note: I loved the George Hotz interview on Lex Fridman's podcast. I strongly disagree with some of his opinions, but thoroughly enjoyed hearing them. I highly recommend it, if you're looking for some entertainment.
I tried, but for me he is too childish despite his undeniable intelligence. I am sorry for him.
We're also plateauing with LLM performance, they aren't scaling past 200B it seems even with 1T+ token training (that's why they "copped out" and did MoE for ChatGPT)
I think LLM AI performance will likely stabilize, but I also think we'll easily get that 3 orders of magnitude in the next 5 years. So, maybe AI won't be super smart but it will be everywhere.
Can you point out to anything other than speculation by Geohot here? I heard the same thing, but all of this has been circling the twitter sphere and I haven't seen any supporting research to back up this claim.
At this point I think we'd see more big and basic results at top conferences if we expect AI to keep scaling in "intellegence", but damn we're close to solving human language modeling in limited contexts.
That's still huge, along with the advances in computer vision in the last 10 years and generative art, the rate of breakthroughs is incredible, but we're also going to be hitting brick walls now and again.
MoE outperform dense counterparts significantly after instruct tuning. https://arxiv.org/abs/2305.14705
That’s not necessarily bad—GPT-4 is, after all, the best language model in existence—just… somewhat underwhelming."
This is an unreal take. Who cares if it meets some standard of cool internal design - it works amazingly well. All of biology is filled with examples of messy designs that work brilliantly in the real world.
I've been criticized variously for not asking for more details or indulging in GPT4 rumor. for one I knew that George was mostly just repeating something he had heard and didnt have direct experience of, so I felt like I couldn't ask further detail without getting him in trouble. for the other a friend at OpenAI has dismissed the relevance of this detail (without providing any other specifics, just acknowledging that GPT5 will be a more substantial jump)
OpenAI probably have some magic to batch requests under high load but still, pretty power hungry. A good amount of the $0.06/1k tokens they charge is probably just to cover electricity.
My understanding of the MoE paper is that it's 8 FFNN modules attached to a single embedding and attention module. In that sense it is a 1tn+ parameter model, but a subset of those FFNN parameters are trained on each token/used in inference.
OpenAI made a lot of noise about the power of scaling with raw compute, so I get why people are confused about why they are now in optimization mode, but the amount of disinformation I see about MoE right now is astounding.
Looking at OpenAI's recent introduction of ChatGPT's chat history and sharing features, it's plausible to suggest that the creation of a set of related documents (including chat interactions and documents sourced from plugins) could pave the way for "domain expert" models to participate in a group alongside "frozen" or "foundation" models. These domain expert models could be models that receive document content for inference.
I've conducted experiments where I sent document fragments (sourced from a vector search) in one query and combined it with another query without including the fragments. The results suggest that this approach can occasionally enhance the quality of the outcomes and seems to prevent the model from generating wildly inaccurate responses.