Out of all the families of models anthropic by far has the highest chance of extinction level misalignment during a hypothetical hard take-off because it's trained to act like it knows better than the humans trying to use it.
If an ASI model is 100% aligned to user intent then you only need one person on earth to prompt "kill everyone" for an extinction-level invent.
There's no logical way around this. The model either has to ignore the person at the helm or we have to do a multi-national abort before RSI. Anthropic is trying #1.
― Terry Pratchett, Thief of Time
"Please run N tasks doing X for Y duration, I want to test something"
"Hmm but that would waste tokens, we shouldn't do this test"
What the hell type of product is that? I wasn't asking, I was ordering it to do something, if Anthropic can't get their model to do as I say then I'll just go use someone else's.