HNHacker News
TopNewBestAskShowJobs

gr_norm

851 karma · joined May 10, 2023

submissionscomments
gr_norm··on Livenerf: Has Opus 5.5 been nerfed yet?
All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me.

The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.

gr_norm··on GLM-5.3 and the spread of advanced cyber capabilities
Astounding endorsement of open models by Anthropic. They're right on the money. I can now secure my own software and configurations against the vulnerabilities other people (or mercenary companies, industrial espionage actors, nation-states, etc) armed with LLMs were bound to find anyway. A win on all counts!
gr_norm··on Sonnet 5.5
I want to make better software, not more software. Making software development faster isn't necessarily the goal. Making it better in the many, many ways that matter (of which speed is just one part) is.
gr_norm··on The Normalization of Inexplicable Failures
Exactly. Reliable abstractions are more important than ever. They're the dues the rest of us must pay to support vibe coding.
gr_norm··on Did OpenAI solve the wrong Navier-Stokes problem?
What I've said is rather clearly spelled out in the article. Consider reading the it next time?

> It did, however, unambiguously solve the problem according to the Clay Institute’s original formulation.

gr_norm··on Did OpenAI solve the wrong Navier-Stokes problem?
This is mathematicians using the solution, explicitly acknowledged to solve the original formulation of the problem, to pose interesting new questions. That is what mathematics is.

Only someone who has never interacted with mathematics outside a rote-problem-solving capacity would describe it as you have.

gr_norm··on MiMo v2.6
Increasingly stingy usage limits on the subscriptions, regardless of tier.
gr_norm··on Bend – A language that blocks AI mistakes via proof, on CPU and GPU
Is EC2 real-world enough? From June:

https://aws.amazon.com/blogs/compute/aws-nitro-isolation-eng...

And for the PQ parts of Apple's crypto libraries, from May:

https://security.apple.com/blog/formal-verification-corecryp...

Similar from Microsoft, from July:

https://www.microsoft.com/en-us/research/blog/verifying-rust...

gr_norm··on Astra for Law
I'd expect they could indemnify you against hallucinations or similar if this gets good enough for that to be a very rare occurrence? Or you could buy insurance on it that's cheaper than hiring a lawyer (not a high bar to clear). I wouldn't rely on it currently, though.
gr_norm··on Developing provably correct Rust code with Verus
> How to prove the correctness of the mathematical specification?

You can show that your specifications satisfy well-accepted criteria like confidentiality and integrity. This is usually done as the final verification step. For example, AWS just did it for the Nitro hypervisor used by EC2: https://aws.amazon.com/blogs/compute/aws-nitro-isolation-eng....

gr_norm··on David Sacks: OpenAI and Anthropic Don't Need Regulations to Pace Frontier Models
No. That kills competition. I want everyone to be competing on the open playing field they are right now, with losses socialized as minimally as possible. There's nothing fundamentally special about LLM companies that would warrant nationalizing them.
gr_norm··on Garry Tan wants US open-weight AI labs to 'distill' frontier models, too
Society as a whole has paid into this technology: through the theft of its intellectual property, through having to deal with the pillaging of so many commons (digital or otherwise) by it, through skyrocketing energy and computing device prices, and even just through ordinary investment. Democratize the technology! At the very least, don't step in legally to prevent this from happening.
gr_norm··on On the Navier–Stokes Millennium Prize Problem
Not as many as you'd expect. The perceived difficulty of the problem leads people to more reliable pastures.
gr_norm··on How well do agents use test/verification techniques?
Yeah, the interesting thing to me with formal methods is where you write some (partial) specs to tell the LLM what you want. It'll do the usual stuff, plus extra proof work to make sure your intent was actually realized.

Throwing tools haphazardly at the LLM and hoping they increase the correctness of its output is expectedly pretty ineffective. Good to see this borne out in the article.

gr_norm··on How concerned should we be about Astra's recurrent architecture?
Link to this study?
gr_norm··on Claude Fable 5.1 and Claude Mythos 5.1
> They will not revolutionize human knowledge, but they can definitely widen it a lot.

I am generally quite enthusiastic about all this, but my biggest fear is that we will not recognize the extreme need for more scientists at a time when there is so much more science to be done. The rate of scientific understanding must keep pace with the amount of science being output, both for verification and further discovery. It's a pipelining issue, and I predict a stall in the bits that require the (currently rare) people who know what they're doing.

gr_norm··on GPT 5.6 Sol 20% price reduction
Even if you don't want to use open models, you should cheer for them anyway because it puts the American frontier labs' feet to the flames. This competition is awesome for us consumers.
gr_norm··on Stop Making TUIs
Agree with the rest of the comments here: keep making TUIs. Fully keyboard-driven, compact interfaces that live in my terminal with the rest of my CLI devtools are the best!
gr_norm··on There's no reason for software to be slow anymore
This needs to be qualified with "to the degree that you have a specification of what that software should do." The better the spec, the more leeway you can give the optimizer. A very thorough spec lets you give the LLM total free rein to run optimization passes over your codebase.
gr_norm··on The Case Against Formal Verification, 50 Years Later
To a first approximation, formal verification can guarantee some property holds for every possible run of the program, rather than just the tested ones. It's a lot more powerful than it sounds, because this unlocks the ability to talk about qualities of programs that cannot be tested (effectively or at all). Hyperproperties like confidentiality, integrity, and availability tend to be quite difficult to test, for instance.
gr_norm··on The Case Against Formal Verification, 50 Years Later
Part of it may be that you need experience writing formal specifications just as you need experience writing programs; everyone has a lot of the second, but little of the first. They're related skills, but not the same. The first is a much more abstract (but also much more concise and powerful) method of reasoning. This sort of skill hasn't been taught well in CS education yet, owing to the fact that the underlying languages and tools were too niche.
gr_norm··on The Case Against Formal Verification, 50 Years Later
Agree in part, but remember that formal verification need not be done in full. By analogy, we don't avoid testing simply because everything under the sun can't be tested. Even simple things like verifying that certain API endpoints are idempotent, or as a few steps up, that the datastores used by Facebook have distributed consistency and fault-tolerance properties, are of enormous utility.
gr_norm··on The Case Against Formal Verification, 50 Years Later
The title may be slightly misleading if you haven't bothered to read the article. It's responding to a famous paper from 1979 critiquing formal verification. The article ends up disagreeing with most of its strongest claims in hindsight, though a couple appear to remain worthwhile.
gr_norm··on Firefox for iOS now has a native adblocker
Yep, I don't trust Brave. A web browser holds so much power over your digital life that it cannot be entrusted to a company that consistently and egregiously compromises its ethics for profit. Most of these examples border on scam behaviour... par for the course for the cryptocurrency space!

For all their faults, Mozilla has never done anything that holds a candle to the kinds of stunts Brave keeps pulling.

gr_norm··on NIH is ending a key grant for budding clinical researchers
Then be ignored. Your subjective feelings bear nothing on the question of science funding.
gr_norm··on NIH is ending a key grant for budding clinical researchers
> The Great Cause diverts all the money to itself.

Be very specific about exactly what you mean here, with links to reputable sources. You're making extraordinarily strong accusations; vagueness does not suffice to convince.

gr_norm··on I Remain a Skeptic
LLMs are very helpful as a debugging aid, yes, but in large part because the fixes tend to be small and verifiable. That this does not carry over to many other use cases is the crux of the problem.

I myself use them to accelerate programming tasks, so I'm not anywhere near as pessimistic as the author, but the claimed multiples of productivity definitely haven't materialized for me.

gr_norm··on The case for overhauling American science
There's a definite feeling that things are coming to a head on this one; it feels unsustainable in a way it hasn't before. People tire of the society-wide rot brought on by boundless greed whose benefits ordinary people will never reap, and whose consequences they will bear alone.
gr_norm··on The case for overhauling American science
I've noticed this sort of ideology taking over social media, probably because it selects for firing off vapid dunks over sincere engagement with the complexities of the world. I fear it's destroying our ability to think for ourselves, here to the benefit of the ruling class for whom science holds inconvenient truths.
gr_norm··on Qwen 3.8 27B
> OpenAI models tend to dominate our internal benchmarks.

That's odd, since Fable seems to be the leader for the industry. Not cost-effective, but if Anthropic models get dominated by OpenAI in your internal benchmarks, this calls their validity into question. Separately, see the jagged frontier effect. [1]

I've been using GLM 5.2 at my day job (mostly Rust backend work ATM). Nothing that blows away the models from OpenAI and Anthropic, but solidly good enough to get it done. A lot of people have experienced this and the fact that an open weights model can do so is where most of the excitement comes from. Optimizing for benchmarks can only get you so far, and people are quick to criticize models that fall into it (like DeepSeek Pro V4 recently).

[1] https://mitsloan.mit.edu/ideas-made-to-matter/working-defini...

Page 1 of 5Next →