This decision is potentially fatal. You need symmetric capability to research and prevent attacks in the first place.
The opposite approach is 'merely' fraught.
They're in a bit of a bind here.
This decision is potentially fatal. You need symmetric capability to research and prevent attacks in the first place.
The opposite approach is 'merely' fraught.
They're in a bit of a bind here.
Cat is out of the bag.
Removing restrictions will help everybody in the long run.
I'm assuming finding vulnerabilities in open source projects is the hard part and what you need the frontier models for. Writing an exploit given a vulnerability can probably be delegated to less scrupulous models.
Good luck trying to do anything about securing your own codebase with 4.7.
I wonder if this means that it will simply refuse to answer certain types of questions, or if they actually trained it to have less knowledge about cyber security. If it's the latter, then it would be worse at finding vulnerabilities in your own code, assuming it is willing to do that.
"This request triggered restrictions on violative cyber content and was blocked under Anthropic's Usage Policy. To learn more, provide feedback, or request an exemption based on how you use Claude, visit our help center: https://support.claude.com/en/articles/8241253-safeguards-wa..."
"stop_reason":"refusal"
To be fair, they do provide a form at https://claude.com/form/cyber-use-case which you can use, and in my case Anthropic actually responded within 24 hours, which I did not expect.
I admit I'm now once bitten twice shy about security testing though.
Opus 4.7 was still 'pausing' (refusing) random things on the web interface when I tested it yesterday, so I'm unable to confirm that the form applies to 4.7 or how narrow the exemptions are or etc.
I once had a car where the engine was more powerful than the brakes. That was one heck of an interesting ride.
So now we have a company that supplies a good chunk of the world's software engineering capability.
They're choosing a global policy that works the same as my fun car. Powerful generative capacity; but gating the corrective capacity behind forms and closed doors.
Anthropic themselves are already predicting big trouble in the near term[1] , but imo they've gone and done the wrong thing.
Pandora is an interesting parable here: Told not to do it, she opens the box anyway, releases the evils, then slams the lid too late and ends up trapping hope inside.
Given their model naming scheme, they should read more Greek Mythos. (and it was actually a jar ;-)
[1] https://thehill.com/policy/technology/5829315-anthropic-myth...