But it's real easy to give auto mode instructions (like "always ask before deploy") and then bypass that just normally.
In one of the occasions it opened a bug report for me just waiting for hit the enter button.
It's easy to "give" instructions, but Claude routinely "forgets" to follow certain instructions, such as "always using the Edit Tool".
Just this week it started to use bash with string concatenation to work around some commands that were blocked in settings.json
The violation above was precisely in this situation :/
My theory is that Anthropic is just a vibe-coding company. Their goal is to capture the attention of white-collar non-coders, since programmers will jump ship fast to another model.
It plays itself.
This is commands blocked in settings.json
I block destructive filesystem operations and destructive git usage via “deny” directives.
I also have instructions injected in CLAUDE.md and re-injected on every single prompt.
Claude just tried to use command concatenation to break those rules. I have also seen it writing a script with rm inside and running it.
For example: our instructions (which are read by the model and classifier) include "do not use sed/python/perl/etc, always use the edit tool for editing", and this only gets followed for a few messages. We have introduced scripts to block those ourselves, since the classifier doesn't care.
Because of those problems, my team is currently testing OpenAI after about a year of Anthropic.