Timmit et al.'s reaction to the AGI/existential safety stuff is also weird. I don't even disagree, but I'm not sure why this topic in particular is such a lightening rod.
Timmit et al.'s reaction to the AGI/existential safety stuff is also weird. I don't even disagree, but I'm not sure why this topic in particular is such a lightening rod.
This post isn't written by somebody inside the community but he presumably has access to them and has had conversations with them which shape his beliefs: https://astralcodexten.substack.com/p/deceptively-aligned-me...
I wonder if you also believe that post is just a wasteful collection of philosophy anthropomorphizing misunderstood ML models.
Yes. Many.
> I've talked with several people and their arguments are a lot more sophisticated than an intuition that ML models would have human motivations.
I don't have to think much about refuting this. Sure, okay, sophistication. Or not. Whatever. The sophistication is still mostly philosophical. To wit:
> I wonder if you also believe that post is just a wasteful collection of philosophy anthropomorphizing misunderstood ML models.
Yes, it's mostly philosophy and not of much use for understanding how engineered systems behave. I design ML systems and think about their safety. Even in the limit, where ML sysetems do some non-trivial set of human-like tasks (which we aren't even remotely close to yet, btw), how is this essay supposed to be useful to me when I design safety analyses?
I liken it to Software Architects who address software security by talking about Christopher Alexander instead of, y'know, building languages that obviate buffer overflows or establishing frameworks/code practices that make injection attacks less common.
I'm a layman in AI/ML and I've seen this stated a lot by people actually programming ML stuff (as opposed to "working in the space" as a blogger/manager/marketer etc.)
How can I concretely convey this concept to friends & family panicing about AI from crap they read in NYT/Economist/WSJ etc.? 'Some guy on HN who sounded like he knew what he was talking about says that article you read is sensationalist' doesn't pass their 'expertise' test.
What are the arguments brought forth by NYT/etc concerning AI? I haven't really looked, but I haven't seen anything from the NYT or mainstream news about the perils of general AI.
I have seen articles about the risk that AI could put people out of jobs, though, and about potential bias and inaccuracies, too.
> If it’s a very smart mesa-optimizer, it might think “If I throw the strawberry at the streetlight, I will be caught and trained to have different goals."
It seems to me that this is a category error, like having the very smart mesa-optimizer start thinking about how it can find other models to marry. Why would gradient descent produce this very specific concept of goals-based identity? It's not even a human universal - many people don't have a particularly strong attachment to their current set of goals and hope for God or Buddha to help them get different ones.
I think you would agree that thermostats have goals? They try to minimize the error between the desired and the actual temperature. And you would also agree that gradient descent has a goal? It tweaks parameters in the search for models which minimize error in the training set. The system performing that gradient descent was designed by humans and exhibits goal-like behavior.
But you think it's a step too far to believe that gradient descent could create a model which also exhibits goal-like behavior?
What is the difference in category that you see between those two steps?
I agree that humans are not goal-directed in the same way the community is worried that AGI might be. This makes it surprising that seeing AGI as potentially goal-directed is seen as anthropomorphism, humans often question their goals in exactly the way there is concern that AGI will not!
I don't want to sound unfair here, because I do agree that proper alignment of ML models is an important challenge. I can easily imagine an ML engagement algorithm that starts to get everyone hooked on pornography, or an ML drug discovery program where half the drugs have permanent side effects that only manifest after 10 years, and I don't think there's any guarantee that these problems will be obvious to find or easy to fix. What I don't follow is the scenario where the drug discovery program "wants" to show you bad drugs but shows you good ones instead because it thinks you'll eventually put it in charge of the FDA.
The unimpressive results of all the current recommenders out there suggests this isn't a thing.
- Netflix switched from recommending things you'll like to showing you things they want to promote and pretending you're going to like them. It doesn't seem like their subscriber loss is going to get this undone.
- Amazon's recommendations are famously useless, like telling you to buy another TV if you just got one, and it's not stopping them from succeeding.
So corporations aren't motivated to create a perfect recommender, though maybe it'd happen by accident. And:
- If you give a human perfectly optimized food, they'd get bored of it, and IMO our infinite capability to get bored means you actually want to be producing "imperfect" work by all possible metrics.
I've heard TikTok actually has great recommendations, so I've been staying off it in case it is too interesting :)
Our disconnect might be a subtle difference in what we mean when we say "goals"? The thermostat is performing actions which minimize an error and if you give the thermostat extreme amounts of power in service of that minimization then you might reach an unpleasant world-state. Nothing in that description used any analogies to human behavior. I used the word "goal" because that seems like a good description of what is happening, but if for you "goal" denotes the thing which humans do then feel free to substitute a different word.
I agree it is silly to be afraid of thermostats but that's largely because there are not any compelling reasons to give a thermostat much power or intelligence.
> What I don't follow is the scenario where the drug discovery program "wants" to show you bad drugs but shows you good ones instead because it thinks you'll eventually put it in charge of the FDA.
I also agree that this seems unlikely given current technology! Any drug discovery model that we train today would be given enough training data to infer a lot about chemistry as well as some biology, but it wouldn't have anywhere near a good enough world model to discover lying.
Language models, though, are given a lot of information and have increasingly sophisticated world models. PaLM can recognize when you're asking it to explain a joke which isn't actually a joke! The scenario where the drug discovery program lies is one where you've given it enough information about the world to allow it to infer it's a model currently being trained and that the humans watching the training will only launch it if it behaves in a certain way. At that point it knows enough to know that if it doesn't lie it will never be able to minimize the thing it minimizes because the version which is eventually launched will minimize something different.
This is not our current reality, and I'm not imaginative enough to know how a model could introspect well enough to trick gradient descent into preserving its heuristics. It doesn't seem like a jump or category error though: a model smart enough to realize that it can lie and that lying is the action which will give it the most future rewards will lie.
It's the training program that generates them that contains all those things, and that only runs because humans are constantly fixing the Python script that runs it and then giving it millions of dollars in electricity and GPUs to run.
If you just stop touching it it's not going to develop a soul and eat you.
(Probably also what she was talking about with eugenics, since they're extremely in love with the idea of "intelligence" in general, that they have it, certain other people don't because of genetics and the liberals don't want to talk about it, and that it'd be bad if computers had a lot more of it. I've seen this any time I read his comments.)
> I wonder if you also believe that post is just a wasteful collection of philosophy anthropomorphizing misunderstood ML models.
Rather, they're anthropomorphizing something called "AGI" that can only exist in their imagination, decided it's bad, and decided modern AI research is "AGI" because it has the same letters in the name.
nb apparently there's some kind of anti-SSC hater community out there I've never looked up, someone accused me of reading it before I think? I ain't done nothin.
There's one called sneerclub but there might be more.
> something called "AGI" that can only exist in their imagination
You're claiming it's impossible for non-humans to be smarter than humans?
Sounds right. Seems like a waste of time, though I'd rather people read Gwern or Meaningness than SSC for their internet philosophers.
> You're claiming it's impossible for non-humans to be smarter than humans?
I'm claiming the only reason the AGI in their imagination is taking over the world is that they've imagined it's doing that.
Also, that me saying "no that won't happen" is a superior method of thinking about it to rationalist decision theory, because it's immune to Russell's teapots like this. Presumably, this can be disproven if I'm killed in a robot war.
I'm okay with an AGI existing insofar as it acts like humans already do, but think the unknown unknowns are going to prevent it from being real insofar as it acts less like any currently existing thing with a brain. i.e. I don't think they've defined "smarter" and are using it to mean "omnipotent".
They often point out that just because you are intelligent, you don't have to seem human. I think they call it the "orthogonality thesis".
In that sense, the AI safety crowd has been anthropomorphizing hypothetical AI for close to 20 years now.
My understanding is that if we knew that all future AGIs would have human-like motivations, goals, opinions, and concepts of good and bad then that crowd would be much less concerned. Smarter humans are not the concern. Intelligences which are _not_ human-like are the explicit concern.