Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
I can sympathize with the argument for the cyber refusals - especially as a temporary measure - especially if Mythos is available to those trying to defend against vulnerabilities.
The LLM development nerfing (and now refusals) is very different though. Anthropic has even said it isn't just for safety reasons:
> Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.
It's at least partially an anti-competitive measure.
The closest analogy is putting measures in a compiler to stop it being able to build other compilers.
Another analogy is priesthoods with secret religious knowledge that "only they are qualified to know".
“The request could assist the development of competing AI models, which is restricted under Anthropic's commercial terms. Benign machine learning work can also trigger this category.”
Source: https://platform.claude.com/docs/en/build-with-claude/refusa...
You’re buying into the hype they’re trying to create here.
Anthropic guardrails seem to be more about protecting their business (distillation), than they are about public safety.
[1] https://www.anthropic.com/news/detecting-and-preventing-dist...
If Google starts calling ads “Best Links” that doesn’t make it correct nor canonical; the correct term is still ads.
Traditionally, distillation is when you get the actual logits of a model response (not exposed via API for years) and then use that to train a model.
Anthropic is offering a commodity product and trying to convince you it isn’t.
It’s even in the name, it’s a myth and a fable. Never happened doesn’t exist.
Also I believe at least on coding that qwen is now the frontier model, fable is its copy of frontier models. In the same way that the Ferrari Luce is an expensive imitation of a SU7 Ultra.
The delusions people live in just to be a hater.
At least the Chinese have the decency of giving back the model weights and not put BS censorship because “it’s too dangerous”.
the problem is so large scale that distill attempts attribute to a decent share of their token revenue generally.
This whole business just keeps getting dumber.
1: https://darioamodei.com/post/policy-on-the-ai-exponential
Frontier AI models, like airplanes, should
be required to go through technical testing
and auditing, and their release should be
blocked or reversed as a threat to public
safety if they do not meet high standards
of safety. I am grateful to see the Trump
administration’s Executive Order move
incrementally towards a greater role for
government in AI, though Anthropic’s proposal
recommends even further action.
They are all-but-literally sucking up to the administration that declared their company a supply-chain risk, arguing that the same administration should be given gatekeeping authority over all high-quality LLMs including open-weight releases. Go gaslight somebody else.First sentence by itself is mundane "regulators are good", which most people agree with, and also libertarians will object to regardless of leader.
Second sentence is obviously sucking up, though is the same level of sucking up found on every stereotypical LinkedIn post.
(Unless they are piping the F1 Mercedes theme song in the announce system at anthropic, in which case maybe you are right)
But I tend to agree, just saying it's a "pretty reasonable statement" and leaving it at that is beyond the pale for anyone who doesn't have an undisclosed stake in the argument.
That is "pretty reasonable" to most people (except the tech-libertarian crowd).
Or maybe he has. I don't know. That would be worse.
I don't really agree with their point here, but there are plenty of people in the AI community whose views are aligned with Anthropic's. That doesn't make them shills.
It's actually important those views are put forward.
A place like LessWrong has the opposite problem - there is no one there who questions the "safety narrative" so the discussion swings more and more towards the extreme end of that spectrum.
My concern is that they won't be pointless in effect. Make no mistake: if Amodei has his way, possession of unvetted model weights will be treated like possession of CSAM is today. And at the same time Amodei calls for that, others are calling for the deployment of technical measures that will make it easier to enforce such laws.
All to the sound of thunderous applause on "Hacker News."
From that paragraph?
Even granting it is sucking up, that is not replacing.
If you think this is OK, I'm not sure what led you to a site called "Hacker News," but fortunately there are plenty of others.
Not sure who you are arguing with really. There seems to be a few logical leaps in between each response. I also didn't say anything like that.
The answer is, the organization making the powerful tool. The people in charge of Anthropic.
Not only that, but they've also written at length about exactly what their opinions and values are: https://darioamodei.com/
You may not agree with the decisions that they make, but they're hardly mysterious. Not something to wonder about.
Again, just because someone has values, doesn't mean they have values you think are good.
What a disgusting country filled with hollow degenerates. No wonder you keep voting for a senile grifting pedophile.
As for making money, you're right, it's not one of your values. It takes a special case of main character syndrome to think it's not anyone's value.
Just before asking for approval to run, it said one thing it wanted to "flag before running" was "Rate-limit and auth testing against prod will generate some 4xx noise in Railway logs and could trip the form rate limiter — harmless, but saying it now."
Ok fine, I said go for it, and it says:
"Running it. Quick recon first (prod URLs + the prior-findings baseline), then I'll fan out the audit tracks with adversarial verification."
Immediately after, I got the Fable warning about how it can't continue because of safety concerns, switching to Opus. In the end, Opus did a good job thanks to whatever Fable suggested doing. Things were fixed that Opus missed in a security/performance audit just the week prior. But what surprised me is that it used 55 agents. Burned 80% of my 5-hour window in 15 minutes (5x Max plan). I've never had Opus do that before on these audits.