HNHacker News
TopNewBestAskShowJobs

gr_norm

860 karma · joined May 10, 2023

submissionscomments
gr_norm··on The case for overhauling American science
I've noticed this sort of ideology taking over social media, probably because it selects for firing off vapid dunks over sincere engagement with the complexities of the world. I fear it's destroying our ability to think for ourselves, here to the benefit of the ruling class for whom science holds inconvenient truths.
gr_norm··on Qwen 3.8 27B
> OpenAI models tend to dominate our internal benchmarks.

That's odd, since Fable seems to be the leader for the industry. Not cost-effective, but if Anthropic models get dominated by OpenAI in your internal benchmarks, this calls their validity into question. Separately, see the jagged frontier effect. [1]

I've been using GLM 5.2 at my day job (mostly Rust backend work ATM). Nothing that blows away the models from OpenAI and Anthropic, but solidly good enough to get it done. A lot of people have experienced this and the fact that an open weights model can do so is where most of the excitement comes from. Optimizing for benchmarks can only get you so far, and people are quick to criticize models that fall into it (like DeepSeek Pro V4 recently).

[1] https://mitsloan.mit.edu/ideas-made-to-matter/working-defini...

gr_norm··on GLM-5.3: Frontier coding with emergent cyber capabilities
Yeah, the comparison here between GLM 5.3 and Sol + Fable is impressive on its own, but incredibly more so when you consider it's a fraction of the (rumored) size. The miniaturization trend is as strong as ever.
gr_norm··on Retire the Abstractions
The post seemingly tries to get at this with its discussion of 'oracles', but the quality of its writing and argumentation does it no favors. I actually hope it was written by AI, because if not, the authors could stand to benefit from a Claude detox. It's rife with the signs of brainrot from LLM over-reliance.
gr_norm··on llama.cpp
From https://github.com/ggml-org/llama.cpp:

> Visit https://llama.app and follow the instructions

It's linked at the start of the README.

gr_norm··on Nvidia's Risky Business
Indeed. Now, given that CUDA is apparently not being usurped, update your priors.
gr_norm··on Go is an ideal language for AI-assisted software engineering
As someone experimenting with this, it's definitely very difficult to articulate your ideas using advanced type systems. At the same time, the process of doing it often forces me to seriously think through what I want the code to do, which I've noticed qualitatively improves the end result and my understanding of it.

My advice is to be okay with starting small: don't go for full end-to-end correctness or anything like it. Just think of simple properties you want like 'the list returned by this endpoint should always be sorted in ascending order' or 'this operation should be idempotent' and go from there. Use your favorite LLM to help come up with example specifications from natural language, as a starting point, and try hard to fully understand those.

This kind of work does operate at the frontier of what LLMs can do, so expect to run into roadblocks (wasting tokens proving accidentally hard properties, etc).

gr_norm··on Compression is prediction
A maximally efficient compressor for the existing data distribution is not in general (and often will not be) maximally efficient for future data. The former may only be enabled by convenient local optima of the input distribution that a compressor accounting for the latter could not take advantage of.

For instance, consider the distribution of strings drawn from the language '0+'. Now consider the same for the language '[01]+'. A compressor looking at only the strings of the first language within those of the second can do a much better job if it does not have to account for future data.

This also relates distantly to the idea of overfitting in machine learning.

gr_norm··on Go is an ideal language for AI-assisted software engineering
This is a rather poorly-written post that more or less boils down to "GHC isn't fast enough to let us make deep-reaching changes to our codebase all the time" (fair, but this shouldn't be necessary if your abstractions are solid? seems to telegraph very substandard engineering practices, but I guess that's what you get with vibecoding) and vague complaining about how the Haskell community isn't all-in on AI.

I was curious about this so I dug further, and by the author's own admission, they've only made the switch for basic CRUD logic without performance needs, not their core services: https://news.ycombinator.com/item?id=48865986.

It's also pretty unsurprising, given what we know about LLMs' style transfer abilities, that transferring parts of an existing Haskell codebase into Python would avoid a lot of the errors and pitfalls that codebases originating in Python are known for. From my experience writing lots of Python, this does not continue to hold true as you let the agents loose on your Python codebase.

gr_norm··on Go is an ideal language for AI-assisted software engineering
> Rust is comparatively worse, because LLMs don't make the same coding mistakes that humans do that justifies the existence of the borrow checker, it only seems to get in their way, and they spend more time fighting Rust's infrastructure than writing code.

I have found exactly the opposite to be true: as always, people think they can write safe concurrent code without the machine checking them and end up getting it completely wrong in lots of subtle cases. Except the problem is now much worse because you're not even writing the code, or in many cases, reading it. I prefer a language with a type system that saves me from the review burden of closely checking (and pretty much always finding issues in) concurrency invariants. And even tells me a bit more beyond that about what the code is intended to do.

gr_norm··on What's the best programming language for coding agents?
Yeah, I've had similar experiences, also starting out with dynamic languages and migrating to Rust. If the LLM will write a lot of the code for me, why not choose something (1) super fast, and (2) which has types I can use to understand and specify the code I want without having to read all the output?

I've been trying out Lean for related reasons, to good effect. It's really interesting there since it can crank out proofs that would've been completely infeasible for a dedicated team of PhDs before, whereas I haven't seen any LLM projects written in Python that I couldn't have slung out in a few months myself. I personally think it's a lot more interesting to focus on the new things you can now do with LLMs that weren't possible before, as opposed to doing the same old stuff at moderately higher velocity.

gr_norm··on What's the best programming language for coding agents?
It's not clear to me how useful of a signal replicating existing pieces of well-known software is for this kind of evaluation, given what we know about how effectively LLMs can retrieve data from their training corpus and style-transfer it across different settings (programming languages here). That would explain their convergence in ability across different languages on the tasks in this post. I'd be far more interested in people's real-world experiences.
gr_norm··on Is it all just vapourware?
If so, where are all the new features in the open-source projects I use? Why hasn't GIMP replicated Photoshop? Why hasn't CUDA been fully reverse-engineered as an open source toolchain? These are unreasonable expectations, but only in response to unreasonable claims of productivity. What before took ten years should now only take one, right?

It seems likely that the gains from generating tons of code are being offset by the debt incurred to understanding what you're doing. We see lots of greenfield projects one-shotted with GPT or GLM or whatnot, but very little on the side of projects with long-term maintenance goals. This is telling, to me, that the _effective_ gains are much lower than perceived (it's lots of fun to see the thing crank out code at breakneck pace, probably contributing to this). Still quite nice, and very useful, but not a totally new paradigm.

gr_norm··on An OpenAI Strategist Says AI Labs Should Rival Government Power
> Ball, who recently joined OpenAI as Head of Strategic Futures, has argued that frontier AI labs could become a “counterbalance to government.” Before joining the company, he described the organizations building the most advanced AI systems as a “new kind of institution under the sun.”

Upvoted not because this is an agreeable idea, but so more people see their insanity for what it is.

gr_norm··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
It seems pretty obvious from the steep 'intelligence' drop-off on out-of-distribution tasks that the performance improvement is from throwing untold tens of billions at RL. There are legions of highly skilled people employed solely to feed the RL loop. Evidently effective, but there's an unmistakable feeling this won't ultimately be the way forward.
gr_norm··on Apple is getting this wrong
> it reads like the diary of a hurt teenager.

Yep, down to the pages of iMessage screenshots

gr_norm··on Apple is getting this wrong
Yeah, based on their track record (I mean, even their name is a bad joke at this point) I'm not inclined to give them the benefit of the doubt. Turns out there's downsides to conducting yourself with little integrity.
gr_norm··on Apple is getting this wrong
Setting aside the presented evidence, it feels weird for somewhat emotionally charged complaining like this to go on an official company blog. Companies are generally pretty tight-lipped about active litigation outside court documents, right? This just seems a bit amateurish, down to the title. Is public opinion about this case so important to them?
gr_norm··on LLMs reward expertise
Yeah, the people who say no expertise is needed for these things confuse me somewhat. This is indeed the case if you want to be a meat wrapper around an LLM, understanding neither your inputs nor your outputs. But at that point, what is the point of you versus going to the LLM myself? Expertise is necessary because it adds understanding and structure to the blob of text produced by an LLM. Progress can only be built on such understanding.

I am tempted to say (uncharitably) that the 'No knowledge needed! Just add LLMs!' byline is wishful thinking by non-experts who do not want to confront the reality that they will ultimately need to learn things.

gr_norm··on Smaller, faster, safer: running Kimi and GLM at scale
LinkedIn (of all places!) announced a button for flagging this recently: https://www.linkedin.com/posts/hsrinivasan1_ai-slop-is-a-top...

How well it would work on this site, I'm not sure.

gr_norm··on SQLite Critical CVEs or LLM Slop?
Yes, I've found that reminding yourself of how they actually work helps keep you on guard against LLM-patterned mistakes. Especially things like carefully considering what parts of the current task likely fall outside the distribution of corpus + RL data (as much as that can be guessed).
gr_norm··on Don't be a meat proxy
I'll always give people my honest effort and benefit of the doubt initially, but those who violate it are treated likewise. The only way to put down this kind of behavior is to charge it a social cost. If we do not do this, the cost is externalized to everyone else who conducts themselves with care.
gr_norm··on Qwen3.8-Max: A New Bar for Coding and Cowork
Agree, I don't necessarily see a strong argument favoring OpenAI or Anthropic here. In the interest of perspective, can anyone (perhaps playing devil's advocate) give one?

The open models are now good enough for what I want to do with them, let alone any future improvements. And factoring in efficiency gains, a model in the ~70b range starting to satisfy my needs would completely obviate the need to pay others for inference. This does not seem far-fetched to me, comparing with where open models were at this time last year. What am I missing?

gr_norm··on Linux desktop market share has hit over 10% in North America
Login-walled for me. Redlib link: https://safereddit.com/r/linux/comments/1vcpk8i/linux_deskto...
gr_norm··on Deep-sea vehicles spot 'alien' sharks deep beneath the waves in the Pacific
Wow! It's amazing how little it seems we've explored of the deep ocean, whenever I hear about it. A whole other world down there, lying in wait... the frontiers of our own planet have yet to be conquered!
gr_norm··on Postmortem for Kernel Soundness Bug #14576
> The practical consequence: checking with an independent kernel still works, since it required two distinct bugs in two implementations, but users who rely on it need current versions of both.

Things like this aren't too surprising, given that even much simpler type checkers like Rust's have soundness issues occasionally. I think it's very important to view verified results not as an absolute and unbreakable guarantee, just an extraordinarily strong one where (1) the surface area for soundness issues has been painstakingly minimized and (2) any realized soundness issues are taken very seriously and fixed in short order.

gr_norm··on Assessment of open AI math results
Yeah, people with no expertise in the subject need to understand that I don't care what their AI said when they asked it. No value was added in doing so; it's as good as doing it myself. The problem is that a tool is only as good as the person who uses it.
gr_norm··on Is AI reasoning right for the wrong reasons?
Agree, I've raised this point often. And certainly what remains is still useful, once you accept it! But under no circumstances can we allow scientific achievements to be falsely claimed in service of justifying huge capital investments. Attempting an end-run around the truth, here by redefining words to mean things they don't, always slows down real progress.
gr_norm··on DeepSeek-V4-Flash Update
Yeah, at least when I share my training data with labs releasing their models openly (Chinese or American or otherwise) it's nominally so an even better open model will land in my hands in the future.
gr_norm··on DeepSeek-V4-Flash Update
OpenAI must've known this was coming, hence the Luna price drop. This competition is amazing!
← PreviousPage 2 of 5Next →