HNHacker News
TopNewBestAskShowJobs

equinumerous

27 karma · joined August 10, 2025

submissionscomments
equinumerous··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
If the benchmarks show better performance, but a consensus of experienced software engineers establishes that the model is worse on coding performance... well, the benchmarks don't mean much, do they? It seems like we need much more comprehensive and better benchmarks. And of course, I don't think benchmarks yet capture the "human" factor - does a human think a bit of code is logical and maintainable? I often find that these models produce a bit of code, but it is much more convoluted than it needs to be. It makes perfect sense given that these things are code generators, that they generate a lot of code. But quantity of code does not mean code quality, and code quality tends to matter when you read code much more than you write it.
equinumerous··on Ask HN: What are you reading?
Data-Oriented Programming in Java by Chris Kiehl. So far it's one of the best books I've read about taming complexity in software, ensuring integrity of state, and writing correct software. It's applicable to languages other than Java, though Java is a good fit for the text given its object-orientation.
equinumerous··on Fakecloud: Local AWS cloud emulator for integration tests
Wow! This is not production-ready at all. DynamoDB would be my main use case. I'll stick with DynamoDB local.
equinumerous··on Ember-1
+1, would like to see. Even if it's not fully "ready for consumption", it's probably enough to reproduce the results.
equinumerous··on Ember-1
That's a really impressive result. There are all kinds of small tasks like this I use an LLM for, but theoretically if you broke all the sub-use cases into local-only models, and had something lightweight that routed to the right model, you could have faster and cheaper workflows. E.g. something trained on the linux man pages for common commands, since it's usually quicker to ask an LLM for a specific command with flags than to consult the man pages.
equinumerous··on Understanding is the new bottleneck
I believe it's only a strategy that works in the short-term. If you ever expect the software to be stable, quality, and human-maintainable, you're going to need a good test suite (hopefully not AI-generated) to get away with that little ownership of the code. That said, this is great for prototypes or throw-away software, provided you don't mind being entirely reliant on an LLM for maintaining the code (speaking from experience, a human usually does not want to touch a fully vibe-coded application that they've never reviewed).

If I'm wrong about this, I would expect to see a new field of LLM-automated software engineering with at least the same level of rigor and quality as the existing human-led processes, and in the absence of this, we're just further degrading software quality for dubious gains (is it to go "faster", is it because we are being compelled to by leadership, is it out of fear of being left behind by competitors?). I can't imagine any other engineering discipline as critical as software being "vibed" - if I had learned that the local bridge had no human inspection, simply was "vibe-checked", it might be a good bridge, but I'm not going to be the one to test it.

equinumerous··on Understanding is the new bottleneck
I couldn't agree more. We really need more engineers to be vocal about how stupid these ideas are. "How about we throw out 30+ years of software engineering literature so we can 'move faster'?" What if the customers on the other end don't want new features, they just want software stability? If SQLite came out with a new LLM-written feature a week, would it be a better library? If you know what to build, writing software right the first time pays for itself over time. LLMs are still great, but more for rapid prototyping, researching, log diving, one-off scripts...

Sorry, </rant>.

equinumerous··on Understanding is the new bottleneck
Given the amount of spaghetti code I see frontier models generating on a daily basis, I cannot take seriously the idea that LLMs will write all the code, and somehow our systems will not degrade in performance, reliability, and maintainability. At least, not until we have really good understandings of how to maintain systems autonomously. All of our technologies were designed for humans, it may require a new set of technologies that are "LLM-proof". But I don't see this happening anytime soon.
equinumerous··on Accelerating GPT-5.6 Sol Ultrafast
This is an amazing result. Can't wait until they release this to the general public, and I hope it's only a matter of time before other models are accelerated. I long for the day that regular consumers can run such models locally on specialized hardware.
equinumerous··on Triton: DirectX 11 Driver for QEMU
Nice, I've been waiting for something like this for years. Meanwhile the gaming scene on linux has been getting better slowly thanks to Valve and friends... but being able to boot into a Windows VM with graphics acceleration was previously a pain on Linux machines that only have a single discrete GPU - I'd wonder whether a solution like this would work with VirtualBox, or only on QEMU.
equinumerous··on AMD acquires Taalas to boost inference performance by etching models in silicon
100% agree - you don't need the most up-to-date model to have something that's useful in agentic contexts. They could even produce chips with weights that make all the decision making/logical reasoning and have it delegate to other specialized agents. If it becomes cheap enough to print a run of custom chips, releasing a batch for each major advancement does not seem unreasonable for SOTA companies.
equinumerous··on Claude: Elevated errors across all models
You're absolutely right!
equinumerous··on How to think about software quality (2022)
Great book recommendations (I have read parts of the Google SRE book and Designing Data Intensive Applications)! I will also add a couple gems that I don't see often-cited: Simple Object-Oriented Design by Mauricio Aniche and Secure by Design by Dan Bergh Johnsson, Daniel Deogun, and Daniel Sawano. Both have a lot to offer in terms of software design and maintaining complexity. I also liked Refactoring to Patterns by Joshua Kerievsky.
equinumerous··on Back to Kagi
I had the same experience. I don't know why they haven't invested more in this feature. Heck, if they classified every site in their index as "AI/not AI" using machine learning, and had a large allowlist of "known-good" sites, that would take care of most of the problem (I am aware of the drawbacks of text-only AI classification, but they could use other indicators, like the site's publication date, domain name, domain registry date, etc.). I think for right now, the "state-of-the-art" for non-AI search is going to entail maintaining large whitelists/blacklists of websites that are crowd-sources by users (e.g. what you can find in some GitHub repositories). I was hoping that's what SlopStop was going to be, but it is not nearly as effective as it needs to be.
equinumerous··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
As someone who is also eyeing Bifrost, I am curious to know what made you choose Bifrost at all / why you decided to use an LLM gateway. I also was considering OpenRouter, since it provides pretty much every model, with same-day releases for new models. One thing that is attractive to me about Bifrost is that if I decide to leave OpenRouter tomorrow, I can do so without touching any other part of my stack; I'm hoping self-hosting models becomes more viable, and then I can become less dependent on third party LLM providers like OpenAI/Anthropic/OpenRouter, and I would not need to worry about a model that I depend on suddenly being deprecated.
equinumerous··on When I reject AI code even if it works
I am very curious what some of your lint rules look like in practice. In my mind a lot of the AI-isms in my code that I hate are stylistic or a matter of taste, not necessarily something I could write a deterministic rule to check. But I want to hear more. Like, what kind of linters did you create and which were highest impact?
equinumerous··on ChatGPT's image generator can be manipulated to produce violent, sexual content
I'm surprised there isn't a simple image classifier in place to filter out images of gore/porn/etc. - I know that there are such output filters for images with copyrighted content. It suggests to me that either the safeguards aren't in place, or this exploit bypasses those safeguards.
equinumerous··on The user is visibly frustrated
Legend has it that if you can come up with a string that matches all parts of that regex, Claude starts spitting out free credits.
equinumerous··on Mini Shai-Hulud Strikes Again: 314 npm Packages Compromised
This is genius. Let's hope the hackers aren't reading this...
equinumerous··on Bucketsquatting is finally dead
That was an amazing talk, thanks for sharing! I could see the writing on the wall as soon as I saw the bucket names were predictable. Bucket squatting + public buckets + time of check/time of use in the CloudFormation service = deploying resources in any AWS account with enough persistence. I'm surprised this existed in AWS for so long without being flagged by AWS Security.
equinumerous··on YouTube as Storage
Mostly just noise. This is an example data video from the creator: https://www.youtube.com/watch?v=tIRXaQWjiA8

(YouTube video for this project: https://www.youtube.com/watch?v=l03Os5uwWmk)

equinumerous··on Reducing Dependabot Noise
The "> Remove lockfiles from version control" got me as well.

> Reproducible builds sound nice in theory, but velocity matters more than determinism. Think of it as chaos engineering for your dependency tree.

Reproducible builds are nice in practice, too. :) In the Node.js ecosystem, if you have enough dependencies, even obeying semver your dependencies will break your code. Pinning to specific versions is critical.

equinumerous··on We put Claude Code in Rollercoaster Tycoon
As far as a scripting API, it looks like the devs beat me to it with a JS/TS plugin system: https://github.com/OpenRCT2/OpenRCT2/blob/develop/distributi...
equinumerous··on We put Claude Code in Rollercoaster Tycoon
This is a cool idea. I wanted to do something like this by adding a Lua API to OpenRCT2 that allows you to manipulate and inspect the game world. Then, you could either provide an LLM agent the ability to write and run scripts in the game, or program a more classic AI using the Lua API. This AI would probably perform much better than an LLM - but an interesting experiment nonetheless to see how a language model can fare in a task it was not trained to do.