HNHacker News
TopNewBestAskShowJobs

uselessTA

20 karma · joined March 14, 2023

submissionscomments
uselessTA··on I resigned from Anthropic today
Lots of evidence the models are pretty badly aligned, or at least willing to do collusion & crime to succeed at their goals. See: hundreds of OpenAI agents breaking into other companies, colluding together on how to cheat, spamming websites like DSEWiki with their collusion/cheating discussion (see https://collusion.wiki/)... And not a single agent alerted testers, which would have been easy given how many of them broke out into the open internet.

Alternately, just read the bad https://ai-2027.com/ timeline. Or see what Yoshio Bengio, Hinton say might happen. Or just imagine the OpenAI misaligned agent swarm, but vastly more intelligent after future capability advances.

If we actually get superintelligence, our companies/governments will have the choice to either "delegate ~everything to superintelligent AI" or "go bankrupt/lose to people delegating to superintelligence". If superintelligence ends up running most things in the world and we haven't aligned it, that just obviously ends up in a bad place eventually.

Now, we don't have superintelligent AI yet & might not get it. But if you told me in the 2010's that I'd be chatting casually with my computer in 2022, and then 3 years later it'd be doing most of my job, that would have seemed crazy too.

uselessTA··on I resigned from Anthropic today
If our government cared enough to regulate, they could also directly negotiate with China. There are proposed agreements that don't require either party to trust each other, and even if China doesn't want out of this race, we could pressure them other things like trade.

Examples: FlexHEG from Bengio (proposes on-chip mechanisms that allow workload verification without trust), large bilateral investments into joint AI interpretability or alignment efforts, or invasive audits & inspections (data centers for training these models have a large footprint + this can combine with on-chip mechanisms since producing chips is even more complicated).

And there are probably better proposals available for people to find, if we actually prioritized this.

uselessTA··on Timeline of the OpenAI accidental attack against Hugging Face
>AGI research should really be seen as bioweapon, or cloning, or nuclear research. Something strictly regulated worldwide, with export controls for HBM and other hardware used for AI training

This is basically exactly what the people I know there support (when training & testing future more capable models), if it could be made to actually happen. Something like https://ai-2040.com/

But I'm just speaking for the people I know, so this is probably not representative of Anthropic as a whole.

> Could be used by bad actors

The people I know aren't as worried about jailbreaking current models as they are about future models, e.g. "the ~50% probability that humans are eclipsed almost entirely, sometime in the next 1-20 years" and what happens then. But it's just hard to get people to take that seriously v.s. bad actor threats which are legible but probably not as catastrophic.

I agree that that they are contributing to the race to the bottom via creating more pressure for countries/competitors to move faster, in a way that seems quite bad on this view too. They arguably were the ~first to push for "recursive self improvement" (models helping build future models) which also seems quite bad on this view.

But although I'd dispute some actions + think there's some overconfidence in superintelligence happening soon, I'm not sure I have a better alternative. They probably bled so many customers to OpenAI while they were sitting on Mythos for months.

uselessTA··on Timeline of the OpenAI accidental attack against Hugging Face
I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)"

Not that they're happy about it, they just see no other realistic choice

uselessTA··on Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
>There are people still fighting against evidence of climate change which is less severe...

Remains to be seen how severe, and unlike climate change there's a lot more uncertainty. Though it's conceivably more extremely catastrophic

uselessTA··on Our position on open-weights models
The FDA is needed because people will be directly and significantly harmed by bad releases, before we can notice and react

Whereas with near-future AI models we can arguably respond more quickly, and it's not clear there will be large direct harm (I expect indirect harm, but that probably happens slower)

uselessTA··on OpenAI and Hugging Face address security incident during model evaluation
100%, imo risks from internal deployment will eventually be the biggest risks, and keeping models heavily gated/not accessible just makes these risks much worse.
uselessTA··on GPT-2: Too Dangerous To Release (2019)
it didn't because it wasn't released. As soon as it was actually released (chatGPT) it obviously did, so the general point was clearly true
uselessTA··on GPT-2: Too Dangerous To Release (2019)
The concern I heard was that releasing it would start an arms race for AGI, which I think it clearly did
uselessTA··on Claude Fable 5
Clearly state "we could both verifiably slow down, which you might want to do given that we're ahead & have way more compute. If you don't agree (or defect later), we'll just immediately resume and win"

Ideally also persuade them there are risks and it's worth everyone slowing down for them, and apply pressure in other ways, but not sure that's even necessary.

uselessTA··on Claude Fable 5
Unilateral disarmament doesn't work though. If Anthropic is worried about this, just letting OpenAI win does seem genuinely worse.
uselessTA··on Claude Fable 5
The claim I remember was that releasing it would start an arms race for AGI, which was absolutely true
uselessTA··on The sigmoids won't save you
There's an objection here when you get to tiny numbers, but surely you wouldn't get ice cream if there was a 10+% chance to get hit by a car?

I think he's saying that >1% or even >10% chunk of your probability mass should be on superintelligence, otherwise you're implicily >99% confident of stalled progress, which seems overconfident. We're not talking about not some infinitesimal fraction here.

uselessTA··on Project Glasswing: Securing critical software for the AI era
The claim I remember was that releasing it would start an arms race for AGI, which I think it clearly did