Here’s something to think about. It’s not an example of a new idea created by GPT like you asked for, but consider what GPT is actually doing. You give it a sequence of text (a prompt), and it is making a prediction of the “best” thing that should come next in that sequence (in the form of a response in this case).
Now consider an arbitrary (but not random) string of binary information. For the sake of example, we’ll use the first 1,000 binary digits of pi. Suppose you have an oracle that returns the shortest prefix-free non-halting universal Turing machine that outputs this string, and you let the machine continue running. What it outputs next is in some sense the maximum likelihood prediction based upon the universal prior (see Solomonoff induction). For our example string, the output will be the subsequent digits of pi starting at binary digit 1,001.
But now let’s consider that instead of starting with a blank tape, the machine starts with a tape containing the information ChatGPT is trained on (presumably large amounts of text sourced from the internet). The oracle now returns the shortest non-halting program corresponding to the conditional Kolmogorov complexity K(user prompt|internet text corpus). (I’m glossing over some nuance related to generating output as part of a conversation rather than as a chatbot monologue).
By replacing the large language model with the oracle, how would you expect the conversation with the chatbot to change? Contrary to what might be assumed, I would not actually expect the bot to appear “smarter” or to generate more novel ideas—despite the oracle—for the simple reason that the prior is merely a corpus of internet text! Why should we expect the maximum a posteriori continuation of this text (+prompt) to contain the solution to quantum gravity anyway? With the oracle, I would expect the conversation to become more “human-readable” in a sense, but not more intelligent.
That said, we can see that ChatGPT is definitely performing inductive inference much better than anything has in the past and possibly at a level that is not even so far from what an oracle would output given the provided prior.
Within the field of algorithmic information theory, it’s well-known that if you can somehow (in the general case) produce a program close in length to the Kolmogorov complexity of the string the program outputs, you aren’t that far away from “near-perfect” reinforcement learning (see concepts like AIXI). And reinforcement learning of that caliber isn’t that far from whatever definition of AGI you wish to use.
So I would argue that ChatGPT is performing inductive inference astonishingly well, but the reason it isn’t “intelligent” is that we haven’t given it the right prior. How do you encode a question like “How can we best cure cancer?” into a prior distribution and a prompt, one for which an oracle would produce the kind of output we are looking for? You first have to make predictions related to intent (what does the human mean, in a technical sense?), then you make predictions related to physics and biology and manufacturing capability (what is the optimal solution to the precise problem, as decoded from an English text prompt?), and only then do you finally layer on a prediction related to encoding the response back into human (what English encoding of the optimal solution will make the most sense to a person?)
To summarize, I suspect the real value of LLMs is their general capability for powerful inductive inference, and once we find the right prior to train the models on (notably, not text from the internet, or even text at all) we’ll start seeing genuine problem solving ability emerge naturally via the inductive ability that is already there.