This seems eerily like the 80s/90s when chess engines were getting smarter, but most people at that time believed they were incapable of truly novel strategies.
This seems eerily like the 80s/90s when chess engines were getting smarter, but most people at that time believed they were incapable of truly novel strategies.
As a practical example - try asking it to make an original joke. It will either fail, or give you a joke, written down word for word, from somewhere else on the internet.
By the time of GPT6 these problems might be overcome, but then it would be more than a language model.
There is plenty of evidence that GPT can indeed create entirely novel ideas, it can do logical inference, imagine scenarios that don't fit its world model, create entirely new pieces of code.
The fact that a Python program is restricted to Python syntax (language model) and libraries (world model), does not prevent it from being Turing complete.
The way I see it, currently, the correct logical inferences made by a model are a by-product of it trying to output acceptable text. That is to say - they are one of the properties, that it has learned, that makes the text acceptable. So it can say that 2+2 is 4 and, of course, even more complex statements like that, but it's based on modelling language, not modelling what is under the language.
> There is plenty of evidence that GPT can indeed create entirely novel ideas, it can do logical inference, imagine scenarios that don't fit its world model, create entirely new code.
I am certainly open to being convinced otherwise, maybe you can give me the best example of a new idea created by GPT?
You should just ask ChatGPT, it will be more convincing than anything I say, I've literally had it make up dozens of things. I had it suggest a niche type system for a novel language to program variational models, and elaborate on top of it given a specific set of constraints. Yesterday I was asking him to make up a new U.I element to replace the timeline in motion graphics software, he proposed a series of bubbles with each bubble representing an effect and its radius representing the length of the effect and changing color when the effect is being played. When it had just come out, I had him imagine the effects of reversing specific principles of lung physiology and how it would affect others. I've had it apply complex functions like ROT13 on entirely original text, encoding numbers in different bases and so on.
If it can execute code on unseen data and perform logical inference on entirely imagined scenarios that don't fit its world model, that's as much evidence of creating a new idea as you can get.
So why did we invent mathematical notation, programming languages, diagrams and all sorts of formal semantics? Clearly it's because natural language does have some characteristics that make it less suitable for some tasks. It may be partly down to our cognitive limitations. But I think the most important issue is the way in which natural language changes. It all happens in a distributed and implicit (on the meta level) way, i.e. without anyone writing a specification of how new expressions relate to meaning.
Having said that, there is no reason to assume that the principles of LLMs are unsuitable to learning formal languages and formal semantics. The question is whether LLMs are the best algorithm for generating and testing formal hypotheses though. They may just turn out to be rather mediocre at this. It's definitely worth trying though and I'm sure this is being done already.
That being said, GPT is also able to generate working code and correct mathematics so there is some evidence he has an internal model of those languages too, and some evidence that those capacities do not strictly require a distinct architecture.
I'm not convinced that it has ideas or that it does logical reasoning. What I can see is that it emits text that is similar to what a person might write if they had ideas or did some (rather flawed) reasoning.
But I also see that it often emits text that a person with these language skills and this breadth and depth of knowledge would never write, because it betrays a complete lack of understanding of what is being said and why it is being said.
For instance, a person is unlikely to casually make up some statistically plausible numbers when asked about a particular company's financials (unless the intention was fraud). That's because a person would have an idea of the purpose of financial reports and the purpose of the question and draw the conclusion that made-up numbers are useless.
And the reason for this strange lapse is not lack of information. If you asked an LLM about fabricated financials, it would go on and on about financial fraud and things like that, suggesting an understanding of the purpose of these numbers as well as very deep knowledge of the subject.
To me, this and many other weird inconsistencies suggest that LLMs have absolutely no idea what they are talking about. And generating statistically plausible sequences of words that were never written before is not the same thing as having a new idea.
In any event, I am optimistic that this technology can be very useful in the right context.
With regards to ChatGPT: Given a set of previously unseen axioms it can (sometimes) make a previously unseen and logically consistent prediction. This is what I have observed.
Whether this is true logical reasoning, truly new ideas, or just an arrangement of statistically plausible characters that belies its lack of understanding is a metaphysical question akin to the Chinese room.
Q : I’ve got a tea cup and a fork. Could you suggest novel ideas about how i could combine these items to come up with a novel use?
A : As an AI language model, I can suggest some ideas:
- Use the fork to create small holes at the bottom of the tea cup, and turn it into a miniature planter for small plants like succulents.
- Use the fork to create patterns or designs on the surface of the tea cup, and turn it into a decorative art piece.
- Use the fork to stir the tea in the cup, and then use it as a paintbrush to create unique and interesting art pieces.
- Use the tea cup as a base for a homemade candle, and use the fork to create patterns or designs on the surface of the candle wax.
- Use the fork to create small holes on the surface of the tea cup and turn it into a lampshade by inserting a light bulb inside.
- Use the fork to create a small opening on the side of the tea cup and turn it into a bird feeder by adding birdseed to the cup.
- Use the fork to create small holes on the bottom of the tea cup and turn it into a soap dish.
- Use the tea cup as a base and the fork as a holder for small items such as earrings or rings.
- Use the tea cup as a pen holder and use the fork to create a unique design on the surface of the cup.
- Use the fork to create small holes on the surface of the tea cup and turn it into a musical instrument by adding beans or rice inside and shaking it.
> No results found for "Yes, but the text space is very sparse."
You see, not even a random phrase of 8 words is duplicated anywhere. Out of all the possible combinations of words that make sense, only a small portion of them actually exist in the real world.
LLMs explore this space that is mostly empty and accomplish useful tasks. They can explore new places by recombining concepts in new ways. They can generate novel ideas by mixing and matching older ideas in ways that make sense.
Now consider an arbitrary (but not random) string of binary information. For the sake of example, we’ll use the first 1,000 binary digits of pi. Suppose you have an oracle that returns the shortest prefix-free non-halting universal Turing machine that outputs this string, and you let the machine continue running. What it outputs next is in some sense the maximum likelihood prediction based upon the universal prior (see Solomonoff induction). For our example string, the output will be the subsequent digits of pi starting at binary digit 1,001.
But now let’s consider that instead of starting with a blank tape, the machine starts with a tape containing the information ChatGPT is trained on (presumably large amounts of text sourced from the internet). The oracle now returns the shortest non-halting program corresponding to the conditional Kolmogorov complexity K(user prompt|internet text corpus). (I’m glossing over some nuance related to generating output as part of a conversation rather than as a chatbot monologue).
By replacing the large language model with the oracle, how would you expect the conversation with the chatbot to change? Contrary to what might be assumed, I would not actually expect the bot to appear “smarter” or to generate more novel ideas—despite the oracle—for the simple reason that the prior is merely a corpus of internet text! Why should we expect the maximum a posteriori continuation of this text (+prompt) to contain the solution to quantum gravity anyway? With the oracle, I would expect the conversation to become more “human-readable” in a sense, but not more intelligent.
That said, we can see that ChatGPT is definitely performing inductive inference much better than anything has in the past and possibly at a level that is not even so far from what an oracle would output given the provided prior.
Within the field of algorithmic information theory, it’s well-known that if you can somehow (in the general case) produce a program close in length to the Kolmogorov complexity of the string the program outputs, you aren’t that far away from “near-perfect” reinforcement learning (see concepts like AIXI). And reinforcement learning of that caliber isn’t that far from whatever definition of AGI you wish to use.
So I would argue that ChatGPT is performing inductive inference astonishingly well, but the reason it isn’t “intelligent” is that we haven’t given it the right prior. How do you encode a question like “How can we best cure cancer?” into a prior distribution and a prompt, one for which an oracle would produce the kind of output we are looking for? You first have to make predictions related to intent (what does the human mean, in a technical sense?), then you make predictions related to physics and biology and manufacturing capability (what is the optimal solution to the precise problem, as decoded from an English text prompt?), and only then do you finally layer on a prediction related to encoding the response back into human (what English encoding of the optimal solution will make the most sense to a person?)
To summarize, I suspect the real value of LLMs is their general capability for powerful inductive inference, and once we find the right prior to train the models on (notably, not text from the internet, or even text at all) we’ll start seeing genuine problem solving ability emerge naturally via the inductive ability that is already there.
Keep in mind that these are first generation systems. If the past 10 years is any indicator, the progress might be unexpectedly fast even for most experts in the field.
These tools may not be independently operational in the physical world yet in 5 years. (Who can say what will happen in 10 years as the arms race is happening now). But if there are humans working with them, their impact can be vast.
LLMs can also learn by reinforcement learning and evolutionary techniques. They don't do just next token prediction, and they can even generate their own training data.
Evolution through Large Models
But it's important to notice that a language model cannot reason to synthesize new information by making this connection. It can only connect two ideas if those ideas were already connected in the training data. To put in another way - ChatGPT is a new powerful way to organize existing information.
And that's not to be dismissive. Natural Language Processing is a hard problem, and ChatGPT gracefully parses through and generates natural language, giving out mostly correct answers at the same time. But the quality of information it gives improves with my skill to ask it good questions. Not different from a internet search engine that gives you better answers if you now how to make better search queries.
You can make a really powerful bomb out of fertilizer but if you aren't a farmer and you place a large order the FBI will star paying attention to you.
Language models predict what to say in a way that makes sense, but a lot of the actual content can be made up nonsense.
It just sounds correct enough to pass for most people.
Thankfully, it has considerable engineering challenges to not be accessible to any individual person or even a small organization. Not only everything is expensive, but it requires decades of hard engineering and construction experience for large-scale projects.
ChatGPT can tell you how to build a nuclear weapon. You can _read_ on how to build a nuclear weapon even without ChatGPT, down to all measures and materials list, this information is not exactly secret anymore. But you on your own won’t be able to do it.
That is the state of LLM's today, I don't think there is a way around this without having some sort of logic engine, similar to how chess engines played and simulated games to train and test things.