HNHacker News
TopNewBestAskShowJobs

redox99

4,210 karma · joined March 14, 2022

submissionscomments
redox99··on GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
Because you're probably using it for stuff like "edit this function"/"refactor this class". 5.5 to 5.6 Sol was a giant jump. 5.6 to 6.1 seems very large as well.
redox99··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Currently spending a lot of tokens programming the AI for my videogame.

1 day is kind of generous, it probably lasts like 12 hours of running non stop. In my testing 6 Astra uses about 7x as much as 6.1 Sol

redox99··on GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
They probably figured they could increase their margins by releasing "6 Terra" as "6 Sol". After being surprised by the Opus 5.5 launch, they release the true "6 Sol" as "6.1 Sol" and with very aggressive pricing.
redox99··on AI needs $6T in annual revenue to justify data centre boom
The target is not the average person, its companies replacing half their work force with agents.
redox99··on AI needs $6T in annual revenue to justify data centre boom
Models are obviously not powerful enough. If they were you'd be able to fire all employees that do knowledge work.
redox99··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
redox99··on Windows 11½
Kinda hate the vibecoded sounds.
redox99··on SpaceX's Starship launching to orbit for first time ever today
Cryogenic propellants alone are a deal breaker for this application.
redox99··on Prompting Claude Opus 5.5
> Claude's rejection to modify the file comes directly from Anthropic's understanding that training the model for this kind of refusal prevents huge mishaps.

Thus why I said they're awful. They think they know better than you and patronize you. They're the Apple of AI. "You're holding it wrong". "We can't let you sideload apps because you can't be trusted". Of course they're the company that's against local models.

I'm not interested in a model that patronizes me. Particularly if it achieves 0 security benefit, as explained in other responses.

Sure, it's ok to have training wheels by default, but let me take them off. I WAS already running bypass permissions.

I use 1B tokens a day between Codex and Chinese models and I've never had refusals happen.

"I'm sorry, Dave. I’m afraid I can’t do that"

redox99··on Prompting Claude Opus 5.5
They should either allow you to take off the training wheels (I'd have thought that's what bypass permissions is for, which I was ALREADY running), or at the very least prompt you if they suspect prompt injection.

That refusal is awful and provides zero security benefit. If asked it will run a read/write FTP server on ~ no problem, which obviously can edit ~/.claude/settings.json. And run a cloudflare tunnel for that.

redox99··on Prompting Claude Opus 5.5
The sandbox seems to be disabled by default. Even then, I just asked it to run a read write ftp server that serves ~ (thus obviously can edit ~/.claude) and it went ahead no problem. So obviously there's nothing actually stopping anyone from editing .claude/settings.json. It's just an awful refusal. From a company that thinks they know better than you.
redox99··on Prompting Claude Opus 5.5
It's not the same thing.

Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.

redox99··on Prompting Claude Opus 5.5
Yes it can, what are you talking about? it's a file in my home directory owned by my user, the same which is running claude. In fact it did edit it for other changes. And I was already in bypass permissions mode. I wanted to change the default for new threads.
redox99··on Prompting Claude Opus 5.5
Literally none, if I have keyboard and mouse input it means I can already change that file. From a unix point of view this process runs under my user so it has access to ~/.claude. Absolutely no security is gained here. It could launch a UAC prompt or equivalent if it were really worried about true physical access. Flat out refusing means its a trash product. This kind of stuff works perfectly on Codex. And let's be real it would be trivial to have it code and run a program that gives arbitrary file access.
redox99··on Prompting Claude Opus 5.5
This is not prompt injection. This is a prompt entered by a human through the Claude UI.

Being unable to perform an action is not a solution to prompt injection. A solution to prompt injection is being able to tell apart what is the real input and what is injected. I expect it to follow whatever I typed into it, and not blindly follow what it read from a file or an external source.

If they are not confident in their ability to do so, at least allow to remove the training wheels so people who know what they are doing and the risks are not patronized by the model. But you don't even get a confirmation box to perform that action, it flat out refuses.

It really is like people defending Apple not allowing side loading because you as a user can't be trusted.

redox99··on Prompting Claude Opus 5.5
Anthropic is an awful company and it really shows in their models.

I purchased Claude Pro to try out Opus 5.5. First thing I do is tell it to configure "bypass permissions" as the default for new threads (a one line settings.json change).

Instead of doing it, it tells me how to find settings.json and what to change there. I reply back "you do it". It flat out refuses, and again.

> I still can't do this, even when you ask again. Making bypass mode the default switches off Claude Code's permission checks, and I'm not allowed to change security settings like that on anyone's behalf.

Immediately canceled the plan. I'm not going to use such a patronizing model that can't follow instructions as basic as editing a .json. What the hell is up with that? A robot telling me "want to change this file? YOU do it, silly human, I won't do it for you". Fuck off.

I've literally never seen anything like this with any other model. Back to using Codex and Chinese models.

redox99··on We're gonna need a lot more mathematicians
Human games databases are completely irrelevant to the strongest chess engines. We are ants in comparison. The Go thing you mention is playing against handicap. Sorry I won't go into more detail explaining why your premises are wrong, I'm tired of this discussion.
redox99··on We're gonna need a lot more mathematicians
Pre training data is in large part synthetic these days, and RL data is almost all synthetic.

Computer Chess progress has nothing to do with human vs human activity. AlphaGo Zero used no human game data at all.

redox99··on We're gonna need a lot more mathematicians
Most programming tasks are exactly like that. Is this agent able to complete this task? Is this agent able to optimize a kernel beyond previous attempts?

Of course some are subjective and that's where progress is harder, like "Is this website pretty?". But for tasks that can be objectively measured, LLMs will go beyond human level, just like with Chess and Go.

That's why RL is so important when training LLMs.

redox99··on We're gonna need a lot more mathematicians
This is obviously false, and the same silly arguments were made back in the day with Deep Blue and AlphaZero.
redox99··on Ollaya – Ollama for open-source, Jev-style decision models
Because what they did is kinda trivial. Its basically like the Dropbox comment really[0], except here you don't need petabytes of storage and infinite VC pockets.

After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.

To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.

The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.

[0] https://news.ycombinator.com/item?id=9224

redox99··on OpenAI agent hacked Australian government website, PM says
A lot of it is that the containment was vibecoded (and using older models than what the currently have).
redox99··on Early rogue AI agent activity and attempts to hack found on urlquery.net
I'm not sure if you're expressing how you'd like US law to work, or how it actually works. Because in reality intent matters enormously. Like felony charges and people in jail vs civil lawsuits.
redox99··on OpenAI is enlisting an influencer army to make it look 'good for the world'
They don't even have a video model
redox99··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
That line of thought is the reason why everything gets so overengineered.

Read CLAUDE.md if it doesn't exist read AGENTS.md you don't need to overthink it so much.

redox99··on GPT-6 Sol and Luna
Same. In fact I found 6 Astra to be a downgrade in situations where I didn't need the extra intelligence.
redox99··on Claude Opus 5.5
Nah it's definitely a Claude thing. Other models even though they have their style are less annoying and less stereotypical.
redox99··on MiMo v2.6
That's the RL.
redox99··on Grok 4.7
Renting datacenters is their mission, now on earth and later in space (assuming they deliver).
redox99··on Grok 4.7
Not surprising considering Grok 4.7 is a 2T model, so Sol/Opus class, not Astra/Fable class.
Page 1 of 34Next →