They're already superhuman in some regards; I don't think that I could have coded up the solution to that problem in 5 seconds. :)
They're already superhuman in some regards; I don't think that I could have coded up the solution to that problem in 5 seconds. :)
> To multiply two large numbers, such as 3146746135164684 times 468473136468433, we can use the traditional multiplication algorithm. This algorithm involves breaking the numbers down into smaller groups of digits, multiplying each group by the corresponding group of digits in the other number, and then adding up the products to find the final result.
> Here is an example of how to multiply 3146746135164684 times 468473136468433 using the traditional multiplication algorithm:
> Break the numbers down into groups of digits. For example, we could group the digits of 3146746135164684 as 314, 6746, 135, 1646, 84 and group the digits of 468473136468433 as 4684, 7313, 6468, 433.
> Multipply each group of digits by the corresponding group of digits in the other number. For example, we would multiply 314 times 4684 to get 1470196, 6746 times 7313 to get 49470158, 135 times 6468 to get 873980, 1646 times 433 to get 708338, and 84 times 1 to get 84.
> Add up the products to find the final result. In this case, the final result would be 1470196 + 49470158 + 873980 + 708338 + 84 = 146845556.
> Therefore, the result of multiplying 3146746135164684 times 468473136468433 using the traditional multiplication algorithm is 146845556.
It's not just that the answer is wrong, is that it's complete nonsense.
Reasoning is a style of thinking that scales. You may be more likely to get the wrong answer in a very long chain of reasoning because at every step you have a nonzero chance of making a mistake, but the mistake is identifiable and explainable. That's why teachers ask you to show your work. Even if you get the answer wrong, they can see at a glance whether you understand the material or not. We can see at a glance that ChatGPT does not understand multiplication.
This is purely ignorant magical thinking.
You're ignoring the very real facts of how ChatGPT works and fabricating a fantasy that you prefer to believe.
In any case, the point is that it's an objective thing anyone can observe: compare the two outputs. See if they differ.
> This is purely ignorant magical thinking.
I counter that you ascribe "magical" thinking to the human condition.
GPT is remarkable to me not for its failings, but because of how far it can get when it's only trying to win the game of "predict what token comes next".
GPT simply gives you a complete hallucination. This can be gotten around imperfectly with prompt engineering to get GPT to admit it doesn't know. But still, hallucinations are pretty dangerous failure modes in many applications, and definitely something we don't want in a future iteration. I don't know how much of a change it is to create a GPT without hallucinations, but I suspect that it is non-trival.
It's a fallacy to anthropomorphize a pile of linear algebra, relate it to a scale of human development, and extrapolate that AI is on a similar trajectory of progress/potential.
Relevant xkcd: https://xkcd.com/605/
If a parrot's brain grew exponentially every year for 20 years, then I might expect it to. :)
> It's a fallacy to anthropomorphize a pile of linear algebra, relate it to a scale of human development, and extrapolate that AI is on a similar trajectory of progress/potential.
I don't believe I said any of those things.
I don't know about parrots, but if you could take a corvid and grant it the same rate of improvement in brain size and training input that these models are receiving, I wouldn't be the least bit surprised to see it doing calculus within 10-20 years.
None invented their own algorithm and confidently claimed it to be the way.
Though they are definitely better at saying I don't know than ChatGPT et al, which have basically been trained to never admit they don't know and always bullshit instead.
How so?
It is a very impressive accomplishment what a large language model can do. It can piece together coherent text from a Google-sized corpus. But I don't thinking describing the process as "reason" is a useful description.
For a subject as purely logical as AI is, it sure seems to draw a in a lot of wishful thinking.
[edit, fwiw]: Also consider that humans can't perceive ultrasonic sound frequencies, or light outside a certain frequency range... We have eventually invented ways to perceive beyond our physical limitations, but no amount of individual adaptation can overcome these boundaries.
There are /other/ arguments that ChatGPT doesn't reason well, but the character-level manipulation examples are insufficient.
Keep in mind that being super confident in a completely wrong answer is way worse than that is.
While all the rocket and space nerds I follow hold Musk in high regard for everything related to SpaceX, all of the civil engineering nerds I follow think TBC is a deadly disaster waiting to happen and that hyperloop is pointless, while all the neuroscience nerds I follow think Neuralink is kinda meh.
If I were to multiply those numbers, it's likely that I'd get the wrong outcome because it's a long computation and at every step I have a nonzero chance of making a mistake. My solution, written out, would look like the correct algorithm, but with computational mistakes along the way. My result could be quite far off, but it would be in roughly the same order of magnitude. If I'd get an answer that's not roughly in the order of magnitude that I expect, I would spot it and -- if so motivated -- start over. If you'd look at my work, you would be able to conclude that I understand multiplication but made a computational error.
ChatGPT also diligently describes its work, and it's just nonsense. The final result is smaller than either of the factors. The algorithm it uses makes no sense. On smaller numbers, the algorithm also doesn't make sense, but it can ballpark the outcome. Therefore, it seems to be mostly using estimation rather than computation to get to the answer, and that estimation breaks down completely for very large numbers.
There's a bunch of famous research that shows a baby and toddlers have basic understanding of physics. If you give a crawling baby a small cliff but make a bridge out of glass, the baby will refuse to cross it, because it's limited understanding prevents it from knowing that the glass is safe to crawl on and it won't fall.
In contrast older humans, even those with a fear of heights, are able to recognize that properly strong glass bridges are perfectly safe, and they won't fall through them just because they can see through them.
What changes when you go from one to the next? Is it just more data fed into the feedback machine, or does the brain build entirely new circuits and pathways and systems to process this more complicated modeling of the world and info it gets?
Everything about machine learning just assumes it's the first, with no actual science to support it, and further claims that neural nets with back-propagation are fully able to model that system, even though we have no idea how the brain corrects errors in it's modeling and a single neuron is WAY more powerful than a small section of a neural network.
These are literally the same mistakes made all the time in the AI field. The field of AI made all these same claims of human levels of intelligence back when the hot new thing was "expert systems" where the plan was, surely if you make enough if/else statements, you can model a human level intelligence. When that proved dumb, we got an AI winter.
There are serious open questions about neural networks and current ML that the community just flat out ignores and handwaves away, usually pretending that they are philosophy questions when they aren't. "Can a giant neural network exactly model what the human brain does" is not a philosophy question.
Not just that -- it's enough to paint a grid on a flat floor using perspective manipulation to make it look like a steep drop.
...on the other hand, I once went onto an open-wall elevator in a VR game, which then started going up. Although I know I had just moved sideways on solid floor and there was still solid floor next to me, it was quite frightening to stand at the edge of it. I had to force myself really hard to tap my foot on the floor outside the elevator that didn't look like it existed.
There's a line of thought that if we're impressed with what we have, if it just gets bigger maybe eventually 'reasoning' will just emerge as a side-effect. This is somewhat unclear and not really a strategy per se. It's kind of like saying Moore's Law will get us to quantum computers. It's not clear that what we want is a mere scale-up of what we have.
> Whether or not they do reasoning, they answer questions with a decent degree of accuracy, and that degree of accuracy is only going up as we feed the models more data.
Kind of. They don't so much "answer" questions as search for stuff. Current models are giant searchable memory banks with fuzzy interpolation. This interpolation gives some synthesis ability for producing "novel" answers but it's still basically searching existing knowledge. Not really "answering" things based on an understanding.
As long as it's right the distinction may not matter. But the danger is a "gut feeling" model will _always_ produce an answer and _always_ sound confident. Because that's what it's trained to do: produce good-sounding stuff. If it happens to be correct, then great. But it's not logical or reasonable currently. And worse, you can't really tell which you're getting just by the output.
> Whether or not they "do actual reasoning" simply won't matter.
Sure it will. There's entire tasks they categorically can't do, or worse can't be trusted with, unless we can introduce reasoning or similar.
> They're already superhuman in some regards; I don't think that I could have coded up the solution to that problem in 5 seconds. :)
This is superhuman in the way that Google Search is. You couldn't search the entire internet that fast either, but you don't think Google Search "feels the true meaning of art" or anything.
Reasoning ability really does seem to emerge from scale:
https://yaofu.notion.site/How-does-GPT-Obtain-its-Ability-Tr...
A recent analysis revealed that training on code might be the reason GPT-3 acquired multi-step reasoning abilities. It doesn't do that without code. So it looks like reasoning is emerging as a side effect of code.
(section 3, long article) https://yaofu.notion.site/How-does-GPT-Obtain-its-Ability-Tr...
Current AIs are fuzzy mad-libs engines. They probabilistically fill in the blank, based on the statistics of all the example text they've seen in training.
In order to predict code successfully, most certainly there's longer-range multi-hop patterns between tokens. (A variable multiple lines away, a variable whose meaning/type/value changes after intermediate lines, ...) Regular english has far less of this.
So it's not terribly surprising to me that including code in training is more effective, relative to not training on it at all.
The question then is say we train on all text ever written in all languages. Including books, code, reddit comments, everything. Could the current architectures and training objectives produce a reasoning AGI? Or are we missing something?
Frankly I think no one really knows for sure yet. I suspect we're missing something and "fill in the blank" is not the key to the universe.
I don't really get this line of reasoning. e.g. I can ask DALL-E to produce, famously, an avocado armchair, or any other number of images which have 0 results on google (or "had" - the armchair got pretty popular afterwards). I can ask ChatGPT, Copilot, etc, to solve problems which have 0 hits on Google. It's pretty obvious to me that these models are not simply "searching" an extremely large knowledge base for an existing answer. Whether they apply "reasoning" or "extremely multidimensional synthesis across hundreds of thousands of existing solutions" is a question of semantics. It's also perhaps a question of philosophy, and an interesting one, but practically it doesn't seem to matter.
If you believe there is some meaningful difference between the two, you'd have to show me how to quantify that.
I don't think you should dismiss it so lightly. This sounds like someone saying the theory of computation doesn't matter...
I see where you're coming from - if it's good enough that we can't distinguish it, then does any difference really matter? I submit it's fundamentally different. This is essentially the Chinese Room thought experiment [1] or a nice similar metaphor with an octopus from Section 4 of this paper [2].
The trouble is not in its ability but human's interpretation it. Humans see an "avocado chair" and think this AI can invent art and concepts. Producing combinations of existing concepts is not "that hard". Even for combinations that have never existed before.
Meanwhile, it's failing at basic tasks: you can find plenty of examples of it failing basic logic, anything with math or arithmetic, a lot of ethics/bias concerns stemming from the training data, etc.
I think when we look forward to an AI that "reasons" this is not what anybody would mean.
Current AIs are bullshit engines. They are very impressive, and probably even useful. They are a milestone. But they are not reasoning in any meaningful way. And if you look at the math behind them there's really no reason to think they would.
So I guess given a methodology that seemingly shouldn't produce reasoning ability, and no evidence that it has so far, sure, maybe scale will magically unlock it. There's always a chance I guess. But it doesn't really seem too sensible.
See Sam Altman's take as well on Twitter here [3]
[1] https://en.wikipedia.org/wiki/Chinese_room
I'm trying to follow your reasoning, but I get stuck right on this line. If you have two systems that are implemented in different ways but the output is indistinguishable I feel that you're forced to claim that the systems operate in fundamentally similar ways. I'm actually confused how anyone could claim the opposite!
I've read about the Chinese Room too and I have roughly the same reaction. I feel like the Chinese Room is akin to saying that any particular neuron in your mind doesn't know how to think. To my view, the guy in the Chinese Room is the same as another one of the many neurons in your head, and it's the system itself which has consciousness, not any constituent part.
So a nuclear power plant and a solar panel farm each generating 4500 MW are operating in fundamentally similar ways?
I mean... in some sense, yeah. They both rely on electromagnetic effects to generate electricity.
They still seem to be operating in really different ways to me, though.
I don't hold with a lot of the evolutionary claims put forward in Jaynes Origin of Consciousness.. but he does put forward a solid intitial discussion of "WTF is intelligence anyway" that's worth the read (if you've not read it and have the interest).
It does. This is trivially true in some domains like mathematics. If you're going to try and measure the last number of pi (which gpt-3 a while ago thought was 3 apparently because some python library returned that result) from observed data I wish you good luck. Deductive reasoning is a necessary skill for any generally intelligent agent. Now how to get from current AI to a system that can reason and how much we're going to need to put in there is an open question, but to deny that deduction is necessary is kind of trivially false. This almost harkens back to the naive empiricism of Skinner that died with the cognitive sciences.
If I assume he is, and his proposed suggestions that the model "participate in a conversation that leads to the kind of questions and answers we discussed here, thereby building trust in the program" and "generate documentation or tests that would build trust in the code" are also in good faith, then I maintain that he's still missing a fundamental limitation of these models even as he outlines its shape with great specificity. They literally and demonstrably are incapable of coherently doing what he wants; they can't be trained to engender trust, only to mimic actions that might by generating novel responses based on patterns.
That would still not be reasoning through the problem to engineer a solution, it's just an extremely effective, superhuman con of novel mimicry. Which, again, is still really, really impressive, and even potentially useful, but in a different way than we might want or expect it to be, and in a dangerous way to use as a stable foundation for iteration toward AGI.
Humans have perceptual systems we can never fully understand for the same reasons no mathematical system can ever be provably consistent and complete. We cannot prove the reliability and accuracy of our perception with our perception.
The only thing which suggests the reliability of our perception is our existence. The better ways of perceiving make a better map of reality that makes persistence more likely. Our ability to manipulate reality and achieve desired outcomes is what distinguishes good perception from bad perception.
If data directed by human perception is fed into these systems, they have an amazing ability to condense and organize accurate/good faith but relatively unstructured knowledge that is entered into them. They are and will remain extremely useful because of that ability.
But they do not have access to reality because they have not been grown from it through evolution. That means that fundamentally they have no error correcting beyond human input. As systems become increasingly unintelligible due to increasing the scale of the data, these systems are going to become more and more disconnected from reality, and less accurate.
Think of how nearly every financial disaster occurs despite increasingly sophisticated economic models that build off of more and more data. As you get more and more abstraction needed to handle more and more data, you get more and more error.
There is a reason biological systems tap out at a certain size, large organizations decay over time, most animals reproduce instead of live forever. Errors in large complex systems are what nature has been fighting for billions of years, and tend to compound in subtle and pernicious ways.
Imagine a world in which AI systems are not fed carefully categorized human data, but are operating in an internet in which 5% is AI data. Then 15%. Then 50%. Then 75%. Then what human data there is gets influenced by AI content and humans doubting reality based categorizations because of social pressure/because AI is perceived to be better. Very soon you get self referential systems of AI data feeding AI and further and further distance from original source perception and categorization. Self referential group think is disastrous enough when only humans are involved. If you add machines which you cannot appeal to and are entirely deferential to statistical majorities, which then become even more entrenched self referential statistical majorities, you very quickly become entirely disconnected from any notion of reality.
How do you know they're accurate if they can't explain how they got the answer?
The problems comes when the data that is fed is of the "Hitler did nothing wrong"-type. That AI system will have no problem regurgitant something that takes that at face value, while a thinking individual knows it to be false.
There's also the issue of what do you do if the data being fed is only "valid" for people that happen to have a certain skin colour? Or a certain ethnicity? Or a certain gender? Or a specific socio-economic status?
There's a great short story about a "robot" ingurgitating lots and lots of data with no extrinsic value in Stanislaw Lem's "The Cyberiad" (minus the Hitler part). People like Norvig are smart enough to give lots and lots of references in order to prove their point but they're not smart enough to see the bigger picture (the one pointed to by people like Lem).
"They are vulnerable to reproducing poor quality training data"
(Peter Norvig, in the article)