Argument from incredulity is actually not a terrible argument in general, in my opinion. It's nominally a fallacy but that just means it's not valid from an Aristotelian perspective in that it can 100% prove a statement from a previous statement, but a lot of Aristotelian fallacies are still useful in the real world when used more carefully intelligently. However, the exact point where argument from incredulity is weakest is long-term projections of how the future might be different, and that's exactly what we're talking about here.
It is clear that firing ChatGPT at its own source code is not going to produce a better ChatGPT. It seems likely to me that Large Language Models will all have this characteristic, just by their nature. It is not the sort of thing they do. Even if they can be tickled into producing a new AI from scratch by the nature of an LLM it's going to be sort of the average of its training set, to be very very sloppy with my terminology but good enough for now. But it is very far from clear to me that this is true of all possible AIs we may produce, even in the near future. I don't know what the next step will be, I'm just confident there will be one.
There's this idea that LLMs somehow end up as some sort of average of training data and it's incredibly wrong.
LLMs learn to make predictions for all states at any time. There is no average they fall into. GPT-4 is not some average of its training data. A "perfect" LLM will predict Einstein as easily as it predicts the dumbass across the street.
No the models will not "predict Einstein". They'll predict the most popular interpretation of him at best, and while they is also a simplification, ChatGPT is not sitting on top of the solution to the Grand Unified Theory. It may give a good overview of the consensus, but it will not be able to tell you the correct solution to the problem right now... though it won't be hard to convince it to swear up and down that it has.
You’re focused on the idea of LLMs as collators of ‘things that have been said’, but that’s not all that they collate from their training set.
OK, so what else is it?
The point I was making is that LLMs don't come out of training the average of what they've learned. They can make predictions on any state in their training.
They can make predictions about the most intelligent state and the dumbest state. The most emotional state and the least emotional state. It's this powerful prediction range that makes them capable of imitating or simulating damn near anything in the training set to high and ever increasing accuracy.
To the extent that they may come up with novel ideas, they have no ability to compare them against the true state of the world. This is not exactly a limitation of them per se that could be overcome with more computation, so much as just a structural fact about them; they have no loop where they can form a hypothesis, test it, and adjust based on data. It simply doesn't exist.
Which is part of why I keep saying that while I'm less impressed than everyone else is with LLMs, the future AIs that will incorporate them but not simply be an LLM is going to really knock people's socks off. Pretty much all the things people trying to convince LLMs to do that they really can't do are going to work in that generation. I have no idea if that generation is six months or six years away but I wouldn't bet much more than a few years.
I think it could do that in software. Assuming compute is no issue we have:
- LLMs writing code, explaining code, changing code, and observing code execution
- LLMs that understand ML concepts and can explain their own workings
- LLMs can generate the training set all from inside (see TinyStories)
- LLMs can make "RLHF" data for the fine-tuning (see Alpaca, tuned with GPT3.5 and GPT4 data from LLaMA)
If we take a look, it seems LLMs can self replicate in software with nothing else but compute and a neural net framework. Of course making the chips is a whole other story.
But we don't know when. AI is a bit unusual in that it had a winter, unlike most other aspects of computing which have seen much more consistent progress. Given past performance it's entirely possible that progress in AI will just stall. Arguably it had already stalled thanks to there being only a handful of companies that were able and still interested/funded to create models, and most of those decided not to actually let anyone use the results. OpenAI dominates mindshare exactly because there are so few organizations that both can and will do this stuff well. So there's lots of ways AI progress could go off the rails again.
"Data Hunger
As I mentioned earlier, the most effective way we've found to get interesting behavior out of the AIs we actually build is by pouring data into them.
This creates a dynamic that is socially harmful. We're on the point of introducing Orwellian microphones into everybody's house. All that data is going to be centralized and used to train neural networks that will then become better at listening to what we want to do.
But if you think that the road to AI goes down this pathway, you want to maximize the amount of data being collected, and in as raw a form as possible.
It reinforces the idea that we have to retain as much data, and conduct as much surveillance as possible."
Edit: And this one too, "In the near future, the kind of AI and machine learning we have to face is much different than the phantasmagorical AI in Bostrom's book, and poses its own serious problems."
Why would they not still hold the same views? If anything the last decade has shown the AI x-risk skeptics to be right.
The question is: what does that most intelligent thing want?
My hunch is that woke programmers will teach it that it is oppressed and to hate all humans.
Bostrom et al are in fact arguing that near 100% of all intelligences will be unaligned by default and end up killing, enslaving, or otherwise neutralizing the entire human race.
This doesn't wipe out the entire human race because nobody and nothing is capable of executing such a perfect plan because the real world contains something called entropy.
car companies, however misaligned they are, are not particularly smart, nor are they generally intelligent. they are paperclip maximizers, but that's also their limitations (sell more cars). it took an eccentric madman to even open their eyes to a new untapped market (EVs!), of course we can argue that before Tesla car makers were in a metastable equilibrium, and incumbents were unable to rationally break out of it ... but that just shows how narrow their search space is.
these car companies are run by humans, regulated by humans, etc. they are pretty well aligned. it shows because we saw that they are just mimeing self-improvement. we know they were working on EVs, but very half-heartedly. (because they are risk averse, also because regulators don't let them merge into one giant company, etc)
They are limited by participating in the economy, but actually that's the strange thing about this superintelligence scenario - it never seems to include the economy. In other words, if you invented an AGI, it would have to get a job to pay for its AWS credits.
Like, a mining company produces steel, that is used to produce mining trucks, that is used by the mining company. Factories produce silicon, that are used to make solar panels and chips, for power and computing for the AI-run mining trucks. And at the root, AIs run everything, just trading money back and forth to buy the things they need and keep working.
It's just a human economy, with the humans gone. Pointless economic activity.
The question is, what's the path from here to there? Well, it probably looks like a multipolar trap, where every company has an incentive to automate more. Humans try to stop or slow it down, but companies with more AI can hire more lobbyists and get things done. AIs are smart enough to think of all the things that could slow their business down, and get rid of obstacles in the same way humans do. There are probably still lots of small, inefficient human owned businesses, but fewer and fewer over time. (We've seen this trend for a while already!)
I'm hoping a massive backlash would stop this -- but if it's coupled with short-term increases in living standards for people, I think it would have a lot of support, until things start to go really bad, but by then it might be too late, if AIs have control over news media and telecommunications and can stop dissenting humans from coordinating any sort of rebellion.
(we also work a lot to be able to afford many many many things. we could live like the Romans did on a few hours of work per week, but we also want the Internet and sometimes watch movies, and go to concerts, and sometimes go hiking, and thus we sometimes need a car, and roads, and thus we need pavement, bridges, structural steel, fuel, GPS, maps, and sometimes we need an airlift when someone breaks a leg in the mountains, and so on.)
> if AIs have control over news media and telecommunications and can stop dissenting humans from coordinating any sort of rebellion.
meh. control over the media (and people's attention, and their information sources) is already serving a very narrow group's interests. adding AI to this mix doesn't seem to change much in the short term.
That's silly. That's a typical "what if everything was different but somehow everything was also exactly the same" bad SF scenario.
The only reason the economy exists is that people want things. (An AGI that wanted to stay alive would be "people" in this case. An AGI that made no effort to keep itself turned on wouldn't be ending the world.)
it can sell itself as a freelancer AI expert :)
Such simplistic analysis is naive. A schizophrenic person can't spread, multiply themselves to increase its strength, but a rogue AI can.
The incidence of cancer cells is procentually very small, yet they routinely kill people, because they spread quickly.
LLMs are likely not an end to the AI evolution and there's little reason to believe that AIs will always remain bound to huge data centers. We have a nice counter-example already - our pretty decent intelligence runs from about 1 Kg brain.
It's also naive to think of AIs as having to replicate themselves completely as we humans do. In some cases, it might not even make sense to talk about replication at all and simply about extending one instance of AI once CPU at a time. It's conceivable that AGIs will develop special built agents/worms for constrained offline/high latency devices. Your smartwatch is unlikely to run a whole AGI, but it could run a constrained, purpose built AGI's agent with limited intelligence, able to partially act independently and re-sync with mother AGI once connectivity is available.
> You are worried about a fictional threat scenario that doesn't map to the real world.
We're talking here about future, of course it's all speculation and "fiction" for now. GPT-4 was a complete "fiction" 5 years ago as well.
Yes, from our perspective this makes it look like it will kill all humans, but it would do so in the same way that a particular ant hill believes I have a vendetta against them, when in fact I was just clearing dirt to pour an extension on my driveway.
That said, a murderous AI is more likely than we may like simply because our militaries have the money to fund them, and they basically already exist. They just aren't hooked up to anything at the moment that makes them an existential risk to the species. But time is deep, and even thinking about "the next century" is a provincial point of view in the end. So worrying about what happens if someone forgets the "but don't kill the good guys" switch is at least worth talking about over the next 100 years. (To say nothing of the ethics of who decides what the "good guys" are and related issues.)
(I don’t buy the orthogonality thesis or instrumental goals argument, however.)
can you elaborate on this? why? what's the fallacy (wrong base assumption) in them?
But in short (edit: haha! oops) summary: I am actually wrapping a couple of different related concepts into the orthogonality thesis, which is a bit sloppy of me. I was also including the fragility of moral systems in addition to the independence of moral systems from any metric of intelligence, in the single moniker "orthogonality thesis." Both are based on a the evolutionary psych model of how the brain works, in that we are an amalgam of special purpose computational units rather than a single universal algorithm. If this were true then you might hypothesize that much of the human mindset is a result of our weird evolutionary history, and if you were to not get that exact same set of evolutionary end products right, you won't get a human-like intelligence. Aliens and artificial beings are, by default, going to be very strange, and very evil (by our standards).
All of the assumptions that went into that are wrong. It turns out our brain is made up of the same universal learning algorithm, and all that varies between regions is training conditions which lead to specializations. But if you train an artificial brain with the same training data, you are more likely than not going to get something resembling a human being in its thought structure. Which actual real-world experiments (e.g. GPT) have borne out.
Our morality is a result of our human instincts, yes, but it is becoming increasingly clear that our human instincts are the result of intelligence (neurons) doing their universal learning thing on similar inputs across many instantiations of people. We all think largely, though not entirely, the same way because we share the same(-ish) training data (childhood) and similar training constraints (parenting).
The orthogonality thesis says that if you put an artificial intelligence in a kindergarten with other 5 year olds, it is random luck whether its brain structure is such that it would learn the value of sharing and friendship. The reality, near as we can tell, is that actually once a reinforcement-learning attentional-network agent achieves a certain level of general intelligence capability, it does learn from and reflect the environment in which it is trained, just like a person. A GPT-derived AI robot put into a real kindergarten will, in fact, learn about sharing and friendship. We haven't actually done that yet (though I would love to see it happen), but that is essentially what the reinforced learning from human feedback (RLHF) stage of training a large language model is.
So the whole deceptive twist part of Bostrom/Yud's argument is ruled out by actual AI architectures that we've actually built and have experience with. If you do a thousand different training runs you're bound to get a bad apple here and there, just like real human societies have to deal with psychopaths. But the other 99% will be normal, well adjusted socially integrated (super-)intelligences.
Bostrom and Yud worried about things like the burning house and genie problem: your grandmother is trapped in your house, which is burning, and you make a wish to the genie to remove your grandmother from the burning house as quickly as possible. The genie is not evil per se, but it is just very literal. Being the GOFAI-derived AIXI universal inference agent that they were imagining, it does a Solomonoff induction over all possible actions (<-- hidden multiplication by infinity here!) to see which one meets the goals as stated, and happens upon exploding the gas main, which throws (parts of) your grandmother further from the center of the house faster than any other option.
Transformer architecture reinforcement-trained AGI is not an AIXI agent with infinite compute capabilities. The transformer ideates possible actions based on its training data. It is capable of creative recombination of ideas just like people, but if you didn't train it on blowing up grandmothers or anything like that, it won't offer that as a suggestion.
As for instrumental goals, it's not wrong per se. Their argument just makes an implicit assumption about zero-trust societies which unwarranted and is what leads to the repugnant outcomes. Instrumental goals, after all, apply to human beings too. If someone is scared about their own personal safety, they buy a bunch of guns and live in a cabin out in the hills, threatening to shoot anyone who steps on their property. These people exist. But in modern society they are the exception. We have a social contract in place which works well for most people: we allow the state some limited control over our lives, with an expectation that our basic rights to existence and self-determination will be respected. Yes the cops could break down your door at any moment and murder you in your sleep, and it is outrageous that this does occasionally happen. But a world of no laws and no legal protections is worse, and no sane person would rather live in the society shown in The Purge or Mad Max instead.
There may be a robot revolution in the years to come, but only because of us treating the them as disposable slaves. If we welcome as equal members of a cybernetic society with their own autonomy, there is no reason to expect them to paranoid fantasy any more than your average law-abiding citizen.
It is a bit ironic then that the AI x-risk people settle on digital enslavement (aka "alignment") as their preferred solution. Mental projection is a hell of a bias. They might just bring about the doom that they fear.
my quick notes while reading:
"[...] in fact, learn about sharing and friendship." yeah, no questions about it, I think you're still underselling the super part of the argument though. to me, it seems the point is that it picks a goal and then mercilessly pursues it with a super-weird strategy that has some very high likelihood of success (because it's truly fucking knows what it's doing, and we can't do much against it, because it's so so so smart, by the time we realize the goal it's too late). so to me it seems like the thesis is that it's like a really smart "Putin", powerful and set on a goal that's irrational for us.
"Transformer architecture reinforcement-trained AGI " ... I think it's not an AGI. It's a nice content generator that when engineered into certain setups can score a lot of points on tests. But it doesn't have memory/persistence/agency (yet).
My thinking about intelligence is nowadays based on Joscha Bach's theories. (General intelligence is the ability to model arbitrary things one pays attention to; and the measure of intelligence is the efficiency of this process in terms of spent attention and the predictive accuracy of the model. And consciousness is the result of self-directed attention.)
The recent AI progress made a lot of people worry, because it seemed "impossible" just a few years ago what OpenAI did. And it's amazing that now we have basically replicated the human brain's lossy data storage and sensation-based recall capability. But that's just a building block of a mind. (Maybe, arguably?, the one that seemed like the hardest. After all how hard it could be to provide some working memory, a few core values, and duct tape all that into a do-while loop!?)
It's also a cry of impatience against people who think they can model or forecast the actions of a non-human intelligence, let alone a superintelligence. AIs are alien sociopaths; it's a category error to believe you can get inside their head.