What these models do is more similar to pastiche, if we really want to compare it to some technique.
But the value of the pastiche depends on the intent and its perception by the viewer(s), because there's nothing inherently original in it.
So basically what models produce has no value, unless we are able to attach some to it. [1]
If we asked chat-gpt to analyze something it has produced, it will probably say that it is "similar to" or "in resemblance of" but it's unlikely that it would say "this is the work of genius, how original! lovely!"
[1] edit: you don't read "pizza maker creates fabulous art on AI" but <person who's already in the business> won a contest of <some art form> submitting something created by <AI of your choosing>. Why? Because they know how to market it and can rely on other people believing that's their creation. Nobody says it upfront "I will submit an AI generated work", because they know it won't be judged the same way. The pizza maker was probably trying it for fun or to make a new logo for the pizza place or simply has no instrument to assign it a value and convince other people that it is true (including using their professional card).
If you write some simple code to generate art procedurally with randomness and it creates really beautiful pieces sometimes, was your software being creative?
I've played around a bit with getting it to write stories etc., and they do often seem quite "creative", the problem is eventually they often end up making little sense or contain fairly obvious contradictions or non-sequiturs in a way I wouldn't expect to see in the output of a typical human author (and certainly not a skilled writer). Indeed I'm not sure ChatGPT has the ability to formulate any sort of longer "story arc" within which to frame the text it generates. But I also suspect it will gradually be able to develop that capacity with future refinements.
When it is wrong it may correct itself, it may double down on being wrong and often just make something up again.
This isn't true. You can ask it basic logic problems that it's never seen before, and it will apply the rules of logic to them. It can also identify correctly which rules of logic would make sense to apply in more complex situations, even when it doesn't get the answer right straight away. At doing this stuff GPT4 is better than GPT 3.5 which is better than previous GPTs. I fully expect that future models will be able to tackle more complex applications of logic to new domains successfully.
If you use only examples that weren't in its training set, you'll get to its limits quickly, but basic level first order logic is definitely within its ability.
As you said, when things are not in its training set it can struggle. If there is a plausible looking text for the question I asked it will give it to me, that's how it is designed. For example I asked it about Windows command line debugger - CDB. It gave me an example command line for it: cdb -c "your-app" -o "logfile". It is very wrong. -c requires an argument which are debugging commands to run on start, -o is to attach to all created attached processes. Real command line looks something like this: cdb -logo "logfile" "your-app" (and it still does not exactly behave as you would imagine having experience with Unix CLI). The problem ChatGPT has with CDB is probably, because it has much much bigger corpus on Unix-like command line tools and because the documentation for CDB is abysmal. From this ChatGPT session I would have more examples.
I'm not saying it is useless. It just is not designed to do that. It might improve or there might be an another algorithm needed on top or instead of what it does use.
For me it is like a kind of a step up from a search engine. It helps me to find something to start with. When I get to some details it is often wrong. I get a starting point from it and then find a proper source for the rest.
For example, it can take code and add types to it. This involves a lot of reasoning ability. It can do this because it’s been trained on a vast amount of code. But it can’t yet fully transfer those reasoning abilities outside the narrow domain of code.
For instance, OpenAI found that GPT4 is much better at reasoning in some human languages than others. It is best at reasoning in English, but struggles reasoning in less resourced languages.
There is clearly some context-independent reasoning going on (i.e. generalization) otherwise the model would not be able to reason at all in languages that it hasn’t seen a particular problem in. But there also appears to be a large context-dependent factor.
In fiction it is often the case that an AI can reason better than humans do, but it doesn't understand emotions. But we now in general that emotions and reading of emotions is simpler than general problem solving. A child picks up on parent's emotions without extensive training. Animals can sense them. The fear response is a basic instinct. I would imagine it should be easier to make a machine being able to almost perfectly recognize emotions than general reasoning or this big heuristic machine which is GPT. I guess it all goes to the training set available.