I'm not talking about AI. I'm taking about intelligence in general. Most Humans would be trashed on a formal logic test. Yes we can be logical no doubt but it's pretty clear we're not running off logic.
If you have an idea of how something works and set off building one from that idea then failing to succeed calls into question the validity of that idea. This is science at its core.
We couldn't really build artificial intelligence the way we thought intelligence worked. So does intelligence really work this way ? That is the crux of the issue.
Intelligent creatures that run off logic basically only exist in fiction. Humans don't do it. Animals don't it. The machines we tried to build to do it this way mostly didn't work. The ones that did work are perfect and never err, more hints humans don't work this way. So we basically have no real indication that logic is at the heart of intelligence.
>Does prediction require an accurate model of the world or not?
Performant Prediction just requires a model, completely accurate or not. Newton's model of gravity is wrong but it's useful fiction so it even still gets taught in schools. Einstein's model is better than Newton's but is still wrong (does not reconcile with quantum mechanics). Doesn't stop it from making useful predictions.
What I'm saying is:
1. Prediction requires a model. The model doesn't have to be "true" to be useful or performant.
2. Perfect Prediction requires a perfect model.
What is a Language Model's goal ? What is the loss trending down to ? It's not to make useful predictions and stop. It is to make perfect predictions of the data they are presented (so of course perfect here doesn't mean a perfect view of the world but rather a perfect view of the world as humans see it)
Everytime the machine uses its existing model to make an erring prediction, it's model is quite literally adjusted and changed to accommodate this error. But by bit this happens.
As a result, what GPT-2 computes is wildly different from what GPT-3 computes and that is different from what GPT-4 computes.
The goal is not to play chess. GPT is not going to stop on some arbitrary line of competency or error rate someone may draw up.
The goal is to model the chess dataset it's been given.
In Evolution, the pressures of two creatures can be just to fly. Plenty of wiggle room.
In gradient descent prediction, the pressures are to fly as some particular creature flies.
I'm not imagining 2 optimizers converging at random. What I'm saying is that the training paradigm is pushing things in that general direction.
>It does bring us back to the original question: if GPT is modeling the world of chess in its answers, why does it make illegal chess moves at a significantly higher rate than humans? Why does it play chess differently from humans?
Is it bad at chess, or... perhaps the way it models chess is fundamentally different from how humans model chess? The simple take is that if you have something hitting 1800 elo but it plays in a way that looks alien to how humans play, it's probably doing something different under the hood.
It's definitely modelling Chess. Othello GPT will construct a board state of the pieces of the game before every prediction. https://thegradient.pub/othello/
But again, models don't have to perfect before they are performant.
The road from A, the nothing predictor to B, the perfect predictor is not instantaneous. GPT is not done modelling Human Chess.
>Those two claims don't jive with each other. If you're arguing that multiple approaches to prediction are possible and that they don't require knowledge or understanding of the underlying systems, then why do you think it's impossible for an AI to make predictions without accurately modeling the world?
I believe I've already explained my point here. When I say multiple approaches, I'm not thinking model(s) vs no model. There will never be performant prediction that does not model its data in some way. But that model can be different. Newton and Einstein paint a very different picture of gravity.
But just like Experimental Science trends towards more and more accurate models of the Universe with time, so does gradient descent prediction trend towards a more and more accurate model of its data.
>If you're arguing that making more and more accurate predictions about the world fundamentally requires an AI to understand the underlying mechanisms of the world, then why are you so laissez-faire about efforts to understand how GPT works when making predictions about its capabilities?
I simply said we can make performant predictions without understanding the internals. I'm being practical. That reality is far more likely than divining the meaning of billions of connections performing computations we had no hand in teaching.
Also recall I said, "As a result, what GPT-2 computes is wildly different from what GPT-3 computes and that is different from what GPT-4 computes."
This means knowing how 3 works internally doesn't mean you know how 4 works internally.