This is just rephrasing "it's not useful for AI." When you say "model intelligence", what you're referring to is AI. I think this is somewhat circular. You're evaluating what techniques are useful for AI. Then you're looking at systems outside of AI and saying "those systems aren't relevant to human learning and can't be how humans learn, because it's not how we build AI."
Do you not see the problem in that logic?
> You don't need to figure out how GPT works internally to make useful predictions on what it will and won't excel at.
The irony of you saying this immediately before you claim:
> It means modeling the underlying structure implicit in the creation of the data. [...] This is what it takes to predict the next token.
Does prediction require an accurate model of the world or not?
I don't claim that you can't work with GPT unless you understand it. My claim is that there's value in understanding it. Understanding the ways it differs from humans will help you work better with GPT. I think that's a fairly uncontroversial claim, there are not a ton of systems where understanding the underlying mechanics doesn't help with performance in some way. It's not essential, but it helps.
If your argument is that modeling the underlying structure that created the data is essential to predicting that data, then it really seems like you should be agreeing with me that understanding the underlying structure that creates GPT output (ie, GPT itself) is important. Honestly, that's a stronger claim than I'm making, I'm just saying it's helpful to know what's going on: you're saying that output prediction inherently requires a model of the systems that created the inputs.
Or are you not saying that? In which case, why are you so confident that GPT is modeling the systems that create its inputs? Is this a requirement or not?
> Humans haven't figured any of that out and we're not doing too bad.
Citation definitely needed, understanding the mechanisms of how humans learn is a huge area of research and would be profoundly useful for us to understand. I mentioned earlier that we decreased reading scores for an entire generation literally because of a bad theory about how humans learn to read. These errors can have large consequences and (in the case of modeling and predicting human behavior) regularly do. That is why there are entire scientific fields dedicated to trying to answer these questions.
----
> On the contrary, that is exactly what ilya sutskever (open ai's lead scientist) believes.
Fair, I didn't realize Ilya was making claims like this. I don't think that's a particularly defensible claim, but fair.
It does bring us back to the original question: if GPT is modeling the world of chess in its answers, why does it make illegal chess moves at a significantly higher rate than humans? Why does it play chess differently from humans?
Is it bad at chess, or... perhaps the way it models chess is fundamentally different from how humans model chess? The simple take is that if you have something hitting 1800 elo but it plays in a way that looks alien to how humans play, it's probably doing something different under the hood.
----
There's some general inconsistency popping up here. Just to emphasize this again: you're simultaneously saying that multiple approaches to task-solving are possible and that we don't need to look at the internals of LLMs or think about how they learn because dumb optimization converges on the same "solution" anyway (the fact that "intelligence" is not a singular solution in the first place is a separate conversation to have). Then you're citing research within LLM communities arguing that human thought patterns mirror LLMs (which I will note, is not the predominant view of most AI researchers) and are arguing that accurate predictive models necessarily must end up accurately modeling and understanding the world in order to make predictions.
Those two claims don't jive with each other. If you're arguing that multiple approaches to prediction are possible and that they don't require knowledge or understanding of the underlying systems, then why do you think it's impossible for an AI to make predictions without accurately modeling the world? If you're arguing that making more and more accurate predictions about the world fundamentally requires an AI to understand the underlying mechanisms of the world, then why are you so laissez-faire about efforts to understand how GPT works when making predictions about its capabilities?
----
My personal take is that we have a system that produces results but produces them in a manner that looks pretty different from what we would expect from humans in the same situation. That probably means that it is using a different method to achieve results. That differing method might affect its performance, and understanding that can be useful for working with the system. For example, you might place a greater emphasis on confirming move legality when using GPT for chess. Or, to be more blunt, you might hook it up to REPL to make sure the code it produces actually compiles because it hallucinates APIs more often than humans do. Or, because you realize that it can't go back and revise previous answers and is creating a predictive output that currently only flows in one direction, you might give it prompts that specifically ask it to break its codegen down into multiple steps instead of fleshing out the entire module. Ie, all of the strategies that are commonly agreed on by pretty much everyone to get GPT to produce better code today; strategies that do not universally map to optimal coding strategies for humans.
I think it's pretty clear that GPT works differently than humans do, and I'm not even certain you disagree with that? You don't seem to be saying that you believe GPT is identically structured to a human brain. One of the differences between GPT and humans appears to be that GPT doesn't always model an underlying reality and think through its answers from first principles. Which again, you seem to agree with! You don't believe that GPT uses first principles when thinking through its answers, right?
Well, apparently controversially, if you give a 7th grader something like an algebra problem, they do use first principles; first principles is how we teach math to kids. It's not the only way that human brains work and it's probably not the core way we understand the world -- we have a lot of stuff happening in our heads; they're complicated and no one really understands human intelligence any more than OpenAI understands what specific algorithms GPT is using for predictions. But (as you yourself have pointed out) we don't teach GPT how to do math the same way we teach a 7th grader, we don't use first principles when teaching GPT. The learning mechanisms and the final strategies are different.