>
Has anyone at OpenAI even made a plausible start at explaining what the hell is going on inside these LLMs that enables them to produce such human-like comprehension?I'm only an amateur in the field, so my uneducated high-level understanding is that, in a sufficiently high-dimensional latent space, there's more than enough dimensions to assign to any single semantic relationship people ever thought of, which is what the training process effectively does, which reduces an important part of thinking - working with concepts and their relationships - entirely to vector adjacency search.
I'm only beginning to study the details, and I don't know how much of specific understanding of this exists, but at the very least this high-level model explains why scaling makes qualitative difference here.
Now, I agree they have strong commercial incentives to push their models as far as possible as fast as possible, but honestly, if I were a researcher working on these models, even if I was somehow unconcerned with any kind of commercial viability and had access to more compute, I'd absolutely keep scaling those models up and up, all the way until I hit the limit of available compute, or the models stop qualitatively improving with scale.
Basically, there's no reason[0] to stop now and try to fully comprehend how GPT-2 works, when GPT-3 was a qualitative jump, and GPT-4 even more so, and GPT-5 is around the corner, and GPT-6 might be a year away from now. All those steps yield important new insights into how the whole architecture works, and if at some point the scaling breaks, that would be even more important knowledge to have. And this doesn't even take into account the fact that, starting with GPT-3, those models are increasingly useful in accelerating both research and scaling alike.
----
[0] - Except, of course, that if transformer models are the road to generic AI, then we'll just blindly race straight into a point of no return.