HNHacker News
TopNewBestAskShowJobs

TuxSH

312 karma · joined February 28, 2020

submissionscomments
TuxSH··on Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026
I doubt it will change much, inference (not training) is very profitable and demand for inference is quite high. See: Mistral serving GLM on their own servers.
TuxSH··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
8x for subscription, 6x for credits/enterprise: https://learn.chatgpt.com/docs/agent-configuration/speed
TuxSH··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
> Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.

TuxSH··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Exactly half as expensive as Opus 5.5 in every API pricing metric
TuxSH··on macOS Golden Gate Is a Buggy Mess
Liquid Glass actually looks like glass now, the weird side panel changes in apps is reverted, music slider hitbox hiccups are now fixed, etc.

It's just 26 but better. On my personal macbook I upgraded from 15 to 27 directly.

TuxSH··on OpenAI: Tomorrow we are re-opening the Pro $200 subscription
Yeah... on OpenRouter or equivalent
TuxSH··on Sonnet 5.5
> OpenAI and Anthropic's lead is vanishingly small at this point.

Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.

Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.

TuxSH··on Claude Opus 5.5
Also Opus 5.5 is often made useless due to its [cyber] guardrails (even worse than Astra), they're even worse than Astra's
TuxSH··on Claude Opus 5.5
One example: https://claude.ai/share/20487190-cf8c-4f25-a7ba-ebfcb1d1a4e9

Notice that this isn't cybersec nor memory-safety related at all.

TuxSH··on I said no and Apple said yes
> I then switched back to MacOS and am mostly happy hanging on to Sequoia… but I’m dreading the point at which (for some reason) Apple make it unavoidable to upgrade.

MacOS 27 is a lot less shit than 26, thankfully. I upgraded from 15 to 27 directly (on a M-series Mac).

TuxSH··on AI coding has made CI a bottleneck, so we reworked ours to keep up
Sure but it's the odd one out in a list of whishing manufacturers added features.

"Switch 2 remains unhacked" is rather disrepectful of the hacking community, and runs on the assumption that the Switch 2 has significant vulns. It's very well possible it doesn't have any high-privilege software exploit at all.

After all if one cas use AI to find vulns and/or to RE, it's even easier for the OS developer (Nintendo) to use it to find bugs in their code before release.

TuxSH··on AI coding has made CI a bottleneck, so we reworked ours to keep up
> Switch 2 remains unhacked

What is Switch 2 security doing here? (independently of it having so few features added per update)

TuxSH··on Grok 4.7
https://knowyourmeme.com/memes/mechahitler-grok
TuxSH··on C++26: Trivial infinite loops are no longer undefined behaviour
FWIW while(true) / for(;;) (or any other loop condition that is a true constant-expression) is NOT UB in C, only C++.

In any case stuff like __asm__ __volatile__("" ::: "memory") prevent such optimizations in the rare case you do need branch-to-self.

TuxSH··on C++26: Trivial infinite loops are no longer undefined behaviour
AFAIK zero-init is unfortunately the default compiler use (but you can change that)
TuxSH··on Claude Code now reads AGENTS.md if there is no Claude.md
> "let's stop AI now"

Pausing AI training would benefit them a lot, as inference is insanely profitable (> 50% margins with maximum demand, afaik)

TuxSH··on ZCode, the GLM coding agent, silently uploads your Git history
Does it do so while using an asymmetric key encryption key?
TuxSH··on DeepSeek v4.1 Flash Is Now Our Best Hacking Model
GLM 5.3 doesn't seem to refuse vuln research (which it classifies as "audit") and is good at it.

Therefore you use for offensive cybersecurity tasks because Daybreak Red/Mythos is pure unobtainium for us mere plebians.

DB Blue thankfully exists, but I suspect you risk a ban if you use it with codebases you neither own nor use

Tl;dr because it's the only model at the level of 5.4~5.6 that doesn't refuse tasks nor risk your oai account getting banned

For plain RE tasks Sol or Astra should work just fine (I think)

TuxSH··on DeepSeek v4.1 Flash Is Now Our Best Hacking Model
Both DS and GLM had the same numbers of subagents, 5 or so.

But, well, number of subagents doesn't make a difference if model is dumb (GPT 5.4 High, in May,in Chat mode outperformed what I see with DS4.1-F).

That being said, pricing model makes a huge difference for "find at least one" tasks: with API/PAYG if you have a chance to save 90%, you go for it, whereas with subscriptions it is optimal to burn all your remaining allowance right before reset

TuxSH··on PS5 Linux lead quits: "a bunch of noobs using LLMs" that "they don't understand"
> If one user could have found this using AI. Then I would imagine anyone else could have found it.

People are still fixated on using AI to produce code (to reduce salary costs/dev time) rather than using it to audit bugs in one own's code, which they are better at.

The latter has "always" been obvious to me.

TuxSH··on DeepSeek v4.1 Flash Is Now Our Best Hacking Model
I find this - or perhaps the title - a bit surprising.

I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.

Perhaps DS works better where targets have low-hanging fruits than can be found fast?

TuxSH··on We must pace the frontier
Of course it is B). If you can use LLMs to find bugs and vulns in software you don't own, you certainly can use them to find bugs in software you _do_ own. They are amazing at that job.

Also, models like GLM 5.3 have zero guardrails in that regard ("find vulns in (...)" prompts just work)

TuxSH··on We must pace the frontier
In terms of API costs they're about right, 10% of weekly usage on a 5x sub is roughly $40 (for Astra). Chinese models cost roughly half as much (and have fewer refusals).

Subs have insane value because large companies cannot use the subscription model. They are loss leaders and are often used for passion projects & startups (where the juicy data is).

Ah and OAI stopped offering the 20x sub as of yesterday.

TuxSH··on DeepSeek v4.1 Flash
> My favourite benchmark for this is to ask it to download a rom for an old game

Even easier: just have them review a large codebase of yours that accidentally has a OOB access bug. Even with no consequences and even if the codebase is truly yours you get blocked.

And of course "find vulnerabilities in..." prompts are out of the question, whereas Chinese models happily oblige.

TuxSH··on Codex on GPT6 Astra Low launched a rocket into space in Factorio Space Age 2.1
Ultra is max reasoning level that splits the work and spawn subagents. Good for whole-project reviews (well, as long as the cybersecurity refusal doesn't appear for whatever reason), but unusable on the $20/mo tier
TuxSH··on Tell HN: OpenAI brings back 5 hour limit for plus and business standard users
It was obvious indeed, but damn does it feel good to have ridden the May 25 - Aug. 26 gravy train.

Now, let's see how the Anthropic IPO goes.

TuxSH··on Show HN: Kadō – open-source habit tracker, with non-binary habit score, for iOS
A README is one of the most important pieces of a project, if one can't even advertise their own project with their own words, it signals low effort.

After all, the people visiting one's repo on GH likely have access to the same AI tools, too.

TuxSH··on Sony makes bold claim about game ownership
On consoles in general, backing up digital media (if even possible) requires full-system exploits, thus bypassing DRM, and decrypting them also requires breaking DRM
TuxSH··on Claude Fable 5.1 and Claude Mythos 5.1
It used to be true up to 2w ago, but with the new/reinstated 5h limits I wouldn't be so sure anymore...
TuxSH··on Claude Fable 5.1 and Claude Mythos 5.1
Unfortunately isn't included in subscriptions and requires usage credits...
Page 1 of 8Next →