When it is wrong it may correct itself, it may double down on being wrong and often just make something up again.
When it is wrong it may correct itself, it may double down on being wrong and often just make something up again.
This isn't true. You can ask it basic logic problems that it's never seen before, and it will apply the rules of logic to them. It can also identify correctly which rules of logic would make sense to apply in more complex situations, even when it doesn't get the answer right straight away. At doing this stuff GPT4 is better than GPT 3.5 which is better than previous GPTs. I fully expect that future models will be able to tackle more complex applications of logic to new domains successfully.
If you use only examples that weren't in its training set, you'll get to its limits quickly, but basic level first order logic is definitely within its ability.
For example, it can take code and add types to it. This involves a lot of reasoning ability. It can do this because it’s been trained on a vast amount of code. But it can’t yet fully transfer those reasoning abilities outside the narrow domain of code.
For instance, OpenAI found that GPT4 is much better at reasoning in some human languages than others. It is best at reasoning in English, but struggles reasoning in less resourced languages.
There is clearly some context-independent reasoning going on (i.e. generalization) otherwise the model would not be able to reason at all in languages that it hasn’t seen a particular problem in. But there also appears to be a large context-dependent factor.
As you said, when things are not in its training set it can struggle. If there is a plausible looking text for the question I asked it will give it to me, that's how it is designed. For example I asked it about Windows command line debugger - CDB. It gave me an example command line for it: cdb -c "your-app" -o "logfile". It is very wrong. -c requires an argument which are debugging commands to run on start, -o is to attach to all created attached processes. Real command line looks something like this: cdb -logo "logfile" "your-app" (and it still does not exactly behave as you would imagine having experience with Unix CLI). The problem ChatGPT has with CDB is probably, because it has much much bigger corpus on Unix-like command line tools and because the documentation for CDB is abysmal. From this ChatGPT session I would have more examples.
I'm not saying it is useless. It just is not designed to do that. It might improve or there might be an another algorithm needed on top or instead of what it does use.
For me it is like a kind of a step up from a search engine. It helps me to find something to start with. When I get to some details it is often wrong. I get a starting point from it and then find a proper source for the rest.
In fiction it is often the case that an AI can reason better than humans do, but it doesn't understand emotions. But we now in general that emotions and reading of emotions is simpler than general problem solving. A child picks up on parent's emotions without extensive training. Animals can sense them. The fear response is a basic instinct. I would imagine it should be easier to make a machine being able to almost perfectly recognize emotions than general reasoning or this big heuristic machine which is GPT. I guess it all goes to the training set available.