Put otherwise, GPT doesn't interact with reality directly, but mediated through our language. But we also don't experience reality directly, so it's hard to see how much of a hindrance that would pose. Hell, much of our model of the world is just as rooted in language as GPT's. The part of my model of reality of which I have direct sensory experience is tiny.
Yes, but the human language is garbage. So while passing throught that filter without maintaining the highest degree of precision barely any of the actual reality remains. Truth sounds exactly the same as a lie.
The actual language that decently captures reality is math. Nothing less is even barely sufficient to produce almost anything but nonsense. Nonsense nicely sounding to human ears but still nonsense.
It certainly "knows" some stuff. However it can't reason about it.
There have been some articles on HN recently on how accurately LMs build world models.
I think the key bit here is that we can influence its associations by inputting the prompt, changing the model that is presented.
I'm waay ahead of myself here, but the thought is interesting, and it will likely remain an open question for at least a few months/years.
Which is actually a novel capability and arises because the network does reinforcement learning over its own context window. It's a strength, not a weakness. Humans can do the same thing. ("Assume that X...")
> It just randomly landed on correct thing first just because it seen it more often in the input data.
Isn't that just a description of learning?
It's true that the network has no idea what is "true". But it's not like we do either, all we do is learning from correlations. We're just better at it.
Humans who can't do math might be a pure language models though.
https://www.youtube.com/watch?v=K0cmmKPklp4
For human students, the median score was 75.6%, ChatGPT performed slightly worse than a Columbia non-science major undergraduate, 73.9%.
Some of the errors was just messing up the arithmetics but it got the reasoning right.
The exam was a reply for this part:
> isn't even close to understanding scientific concepts and knowing how to connect them
That test in my book definitely proves that it can understand (or let's say process and create abstractions) scientific concepts and can place them on a knowledge graph.
> So what is really the utility of "testing" it?
Eventually these models will be integrated into everyday products so we need to learn how trustworthy they are. These tests are written in natural language, the same way everybody interacts with them.
If you really do think this stuff is important, show it on its own terms. Make tests suited for the model with rigorous expectations tuned for a computer, not a human. Then you can start showing some impressive things. I just urge you to see the basic flaw in your conceptions here. I don't want to rob you of your enthusiasm, just better direct it.
Nobody is saying that GPT-3 is smart. But it is surprisingly capable.
> Test's are useful precisely to gauge human retention/understanding
Indeed, and these models are also compressing information they were trained on. It makes sense to test them how well can we access and query that knowledge.
I'm not saying at all, that passing tests make them human or anything like us, but it speaks a lot about the quality of the model. It's not the ultimate quality check of course, but it's an important one, because you can assign a score to the results of the test.
> Its like applauding the calculator for doing math.
Show that calculator to Charles Babbage, he would be in awe :) This is probably the leap why most people are freaking out, it feels a significant step towards intelligent machines regardless of the system under the hood.
I actually tested ChatGPT, (like a few million other people). I'm not saying it was done with scientific rigour of course. It passed a few flawlessly and failed some other. I believe I know what I'm interacting with, I've read about how GPT works and I'm planning to learn more. Don't worry about my enthusiasm :)
> 7:36 The exam is structured with multiple choice questions at the start and then some freeform short answers towards the end.
Imagine if you could look up all material in the last 10 years published about astrophysics (or insert any major topic here) near instantly, the question's details are also presented for ingestion, then make prediction what is and is not based on weighing the frequency and potentially the source of information, and then respond in relatively coherent sentences.
There is no astrophysics in this.
That's why, if you want ChatGPT to actually think about a problem, you have to get it to give the explanation first, gradually, and only then output the answer. That way, GPT can use the rationale it has given in the prediction of the answer that fits best. If it has to give the answer first, it's already committed.
If a human has to answer a question, they follow a process like this:
- parse the question
- compute a rationale, and evaluate it to gain an answer
- output the answer first, because it's the most useful
- output the cached rationale.
But GPT trains left to right, and it cannot see the first two steps as they're not in the text. From its perspective, most humans produce answers and explanations simultaneously, instantaneously and intuitively. This makes it almost impossible for GPT to learn how to think. It has to rely on the very few samples where humans give the rationale first, and then the answer. And because this is rare, you have to really push GPT to actually use its scant cognitive abilities. ("Let's think about it step by step.")