312 karma · joined February 28, 2020
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
It's just 26 but better. On my personal macbook I upgraded from 15 to 27 directly.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
Notice that this isn't cybersec nor memory-safety related at all.
MacOS 27 is a lot less shit than 26, thankfully. I upgraded from 15 to 27 directly (on a M-series Mac).
"Switch 2 remains unhacked" is rather disrepectful of the hacking community, and runs on the assumption that the Switch 2 has significant vulns. It's very well possible it doesn't have any high-privilege software exploit at all.
After all if one cas use AI to find vulns and/or to RE, it's even easier for the OS developer (Nintendo) to use it to find bugs in their code before release.
What is Switch 2 security doing here? (independently of it having so few features added per update)
In any case stuff like __asm__ __volatile__("" ::: "memory") prevent such optimizations in the rare case you do need branch-to-self.
Pausing AI training would benefit them a lot, as inference is insanely profitable (> 50% margins with maximum demand, afaik)
Therefore you use for offensive cybersecurity tasks because Daybreak Red/Mythos is pure unobtainium for us mere plebians.
DB Blue thankfully exists, but I suspect you risk a ban if you use it with codebases you neither own nor use
Tl;dr because it's the only model at the level of 5.4~5.6 that doesn't refuse tasks nor risk your oai account getting banned
For plain RE tasks Sol or Astra should work just fine (I think)
But, well, number of subagents doesn't make a difference if model is dumb (GPT 5.4 High, in May,in Chat mode outperformed what I see with DS4.1-F).
That being said, pricing model makes a huge difference for "find at least one" tasks: with API/PAYG if you have a chance to save 90%, you go for it, whereas with subscriptions it is optimal to burn all your remaining allowance right before reset
People are still fixated on using AI to produce code (to reduce salary costs/dev time) rather than using it to audit bugs in one own's code, which they are better at.
The latter has "always" been obvious to me.
I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.
Perhaps DS works better where targets have low-hanging fruits than can be found fast?
Also, models like GLM 5.3 have zero guardrails in that regard ("find vulns in (...)" prompts just work)
Subs have insane value because large companies cannot use the subscription model. They are loss leaders and are often used for passion projects & startups (where the juicy data is).
Ah and OAI stopped offering the 20x sub as of yesterday.
Even easier: just have them review a large codebase of yours that accidentally has a OOB access bug. Even with no consequences and even if the codebase is truly yours you get blocked.
And of course "find vulnerabilities in..." prompts are out of the question, whereas Chinese models happily oblige.
Now, let's see how the Anthropic IPO goes.
After all, the people visiting one's repo on GH likely have access to the same AI tools, too.