4,210 karma · joined March 14, 2022
1 day is kind of generous, it probably lasts like 12 hours of running non stop. In my testing 6 Astra uses about 7x as much as 6.1 Sol
Thus why I said they're awful. They think they know better than you and patronize you. They're the Apple of AI. "You're holding it wrong". "We can't let you sideload apps because you can't be trusted". Of course they're the company that's against local models.
I'm not interested in a model that patronizes me. Particularly if it achieves 0 security benefit, as explained in other responses.
Sure, it's ok to have training wheels by default, but let me take them off. I WAS already running bypass permissions.
I use 1B tokens a day between Codex and Chinese models and I've never had refusals happen.
"I'm sorry, Dave. I’m afraid I can’t do that"
That refusal is awful and provides zero security benefit. If asked it will run a read/write FTP server on ~ no problem, which obviously can edit ~/.claude/settings.json. And run a cloudflare tunnel for that.
Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.
Being unable to perform an action is not a solution to prompt injection. A solution to prompt injection is being able to tell apart what is the real input and what is injected. I expect it to follow whatever I typed into it, and not blindly follow what it read from a file or an external source.
If they are not confident in their ability to do so, at least allow to remove the training wheels so people who know what they are doing and the risks are not patronized by the model. But you don't even get a confirmation box to perform that action, it flat out refuses.
It really is like people defending Apple not allowing side loading because you as a user can't be trusted.
I purchased Claude Pro to try out Opus 5.5. First thing I do is tell it to configure "bypass permissions" as the default for new threads (a one line settings.json change).
Instead of doing it, it tells me how to find settings.json and what to change there. I reply back "you do it". It flat out refuses, and again.
> I still can't do this, even when you ask again. Making bypass mode the default switches off Claude Code's permission checks, and I'm not allowed to change security settings like that on anyone's behalf.
Immediately canceled the plan. I'm not going to use such a patronizing model that can't follow instructions as basic as editing a .json. What the hell is up with that? A robot telling me "want to change this file? YOU do it, silly human, I won't do it for you". Fuck off.
I've literally never seen anything like this with any other model. Back to using Codex and Chinese models.
Computer Chess progress has nothing to do with human vs human activity. AlphaGo Zero used no human game data at all.
Of course some are subjective and that's where progress is harder, like "Is this website pretty?". But for tasks that can be objectively measured, LLMs will go beyond human level, just like with Chess and Go.
That's why RL is so important when training LLMs.
After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.
To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.
The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.
Read CLAUDE.md if it doesn't exist read AGENTS.md you don't need to overthink it so much.