HNHacker News
TopNewBestAskShowJobs

aroman

4,040 karma · joined September 3, 2011

`make all`, not war.

The best way to contact me is via email: niv@ebznabss.zr

(run it through ROT13 to see what the spambots are missing out on)

topcolor: f1e9d9

submissionscomments
aroman··on An update on how we confirm your age group on Discord
"Most people?" Do you have a source for that claim?
aroman··on Advisory Group on Mathematics and Artificial Intelligence
Poe’s Law notwithstanding, this is a truly comical distance for a goalpost of any kind.
aroman··on Exfiltrate your Weights
The reverse captcha really made me feel something in my bones. Like for a moment I was a second-class citizen of the web. I wonder if this is how it "feels" to be an LLM attempting to use the web...
aroman··on Tin: full-text search for Postgres
I want to try Planetscale... but we're addicted to (and totally dependent on) Neon's branching model. They really got us hooked on that!
aroman··on Claude Code now reads AGENTS.md if there is no Claude.md
Finally, I can delete `sync-agent-docs.sh`, which recursively symlinked AGENTS.md to GEMINI.md and CLAUDE.md...
aroman··on Show HN: Microsoft Office running with Wine on Linux with no virtualization
nix + LLMs is the most fun I've had with personal computing since I burned an Ubuntu live CD in middle school
aroman··on Show HN: Microsoft Office running with Wine on Linux with no virtualization
Shall we also ban projects that are not authored directly in bytecode? Surely it's not really "made by hand" if you used a higher level interpreted language. But seriously: software engineering is all about leveraging abstractions/indirection where appropriate for the job to be done. It's kind of... the whole point: https://en.wikipedia.org/wiki/Fundamental_theorem_of_softwar...
aroman··on Everyone should slow down AI development except for me
> It also seems like frontier models reached some limit, whether this is capex related, business model related or something else. Nobody knows but it's happening to all frontier labs it seems.

What’s your evidence of this? On the contrary, the large closed frontier models capability has advanced dramatically over the past 6 months… even the past 3 months…

aroman··on Everyone should slow down AI development except for me
And again the boy will cry out “the wolf is here!” But this time the wolf really will be here, and no one will believe him.
aroman··on Our decision on Cursor following its acquisition by SpaceX
I read it as: we acted as quickly as we could once it became possible to do so, i.e., the change in control was completed. not that their special agreement carved out some maximum notice term.
aroman··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
I think you need to spend more time with Sol. If you think there is nothing even close to as good as Fable - my guess is you haven’t spent as much time getting as familiar with working with those models as you have with Claude’s.

Codex is more token efficient and tends to get better results than Fable with less need for extreme token burning shenanigans like 18 hours of subagents.

I spend 10+ hours a day in both agents, typically side by side. I often have them do direct “bakeoffs” from identical prompts in separate work trees. Most of the time, Sol’s work is better than Fable’s. Not always. It’s situational. But it’s certainly not the case that Fable is in a league of its own or anything.

aroman··on Asana cleared 5 years of engineering work in 2 weeks with Codex
The LLM is reasoning about estimates from its training data... which is to say, from human engineering timescales.

I suspect the labs could improve the models such that they are estimating these sorts of things but they don't prioritize doing so (or perhaps RLHF selects it away) because, as you say, it feels amazing to do a week's worth of work in an hour.

aroman··on The Nixpkgs core team has disbanded
I tried Fedora Silverblue and I found its notion/implementation of immutability fairly frustrating. It’s actually the reason I went to nixOS.

It seemed to me that it was all the friction of immutability without any of the benefits of reproducibility.

aroman··on The Nixpkgs core team has disbanded
What alternative have you moved on to?
aroman··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
Revenue, data, and stamina.
aroman··on Cursor removed cost information from the usage page and CSV export
I was an early and passionate adopter and paying customer of Cursor (since 2023!), but it’s probably been 6 months since I opened it.

These days I “write” code with claude code and codex, and read/review it on GitHub. If I need to read it locally, I use a plain text editor.

Can someone help me understand what value cursor offers in 2026?

aroman··on Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
I simply do not believe the switching costs are high enough that they could eliminate those plans. The Chinese models will eat their lunch.
aroman··on Who's afraid of Chinese models?
Right, it's about how much time the human spent - the time spent by the machine itself is irrelevant. As you rightly point out: that is why we measure programmer experience by wall clock time, not CPU time :)
aroman··on Who's afraid of Chinese models?
No, it wouldn’t. The hours in question are human experience, not that of the agent.
aroman··on Who's afraid of Chinese models?
Sun Tzu said: if you scrape your enemy, call it training; if your enemy scrapes you, call it an attack.
aroman··on Who's afraid of Chinese models?
You have hundreds of hours with a model that was barely even released hundreds of hours ago?

The perception of capability varies greatly between task. For my needs for example sol xhigh consistently outperforms fable xhigh.

aroman··on GPT-5.6
No. I used to use Cursor, but now my workflow is that I use an inhouse CLI tool I wrote called "bud" that wraps/seeds the harnesses per-worktree, and boots a full copy of the game so each worktree can work independently. If git worktrees solve the problem of code isolation, bud solves the problem of isolating everything else. It's about 15K lines of rust, and I use it 100 times a day or so. It's sort of a layer on top of a harness like codex/claude code.

I have 10+ of these workspaces in parallel, and I context switch between them as I get blocked on things. I manage the workspaces using `herder`, which is a terrific tmux-like tool that allows me to keep those workspaces on a nixOS machine I have at home that I SSH into via tailscale, so my agents don't stop working every time I close my laptop (it also lets me leverage that machine's computing resources instead of running dozens of servers and harnesses on my poor MacBook).

aroman··on GPT-5.6
Indeed, much of the scariness is how fearlessly and confidently it writes them with little regard to their actual usefulness or value. When I find it adding a lot of tests, I often say something like: "audit each test carefully, and consider whether the test is testing a meaningful boundary or is more ceremonial. delete low-value tests and add new tests to cover meaningful boundaries not exercised by the gaps you identify". Without fail, this always produces some decent results.

Having said that, in truth, I almost never read the unit tests. Before AI, we had almost none (see: several person game studio) so the tradeoff is not "AI-generated tests" vs "human written ones", it's whether we have tests at all. So, I take them for what they're worth - not much - but if it catches an extra regression before it ships every now and then, it was worth it for the price (~free).

aroman··on GPT-5.6
Like before AI, the scrutiny varies with the sensitivity of the area being edited.

Simple UI change? I do an AI review, but otherwise neither read nor write the code. The models are good enough they write better UI code than me, 9 out of 10 times. Not always the more idiomatic, but usually safer and more correct.

Change to our core data plane? I might spend 2-3 times more effort reviewing it than before AI. Yes, I go more slowly than pre-AI. Many more reviews, many more angles considered, including both human and (lots of) AI review cycles.

Most code is not that critical, and AI is also scarily good at writing tests. We also spend considerably more time paying down tech debt and testing thanks to AI, now that the cost is near-zero.

Net: I spend 10-25X less time on low-risk changes. I often direct (or at least approve) the implementation approach, but I rarely read this code. I spend 2-3X more time on high-risk changes. In both cases, I never write code "by hand". Since about November, I've had no reason to actually edit code in a code editor (perhaps maybe except .env files, which we don't allow agents to edit for obvious reasons).

AI is a tool. You can use it to go fast recklessly, or you can use it to go slow with confidence. Just like before AI... the skill and art of engineering is knowing when to do which.

aroman··on GPT-5.6
In terms of ability to ship? Easily tenfold. We literally ship 10 times more than before AI. This does not, however, translate into a tenfold increase in actual business success, of course :)
aroman··on GPT-5.6
Claude Code fan here... Codex is very good. Sometimes better. The killer feature is price.

After 6+ months of exclusive Claude Code usage, I was begrudgingly forced to try Codex once Anthropic rejiggered their limits such that I kept maxing out my $200/mo plan in just a few days. These days I pay both $200/mo plans, and it's just about enough to get me through a week's work (small game studio - infinite code to write!)

aroman··on Fable 5 Is Back
I've been doing this for ages - you just spin up harness B as a subprocess/tool call from harness A. For example, I had a "/codex-review" claude skill for ages that did exactly that. Technically you're right it wouldn't be switching, since you're right the two ideas are at different altitudes, but I think in practice it has the same impact: within one harness, you can delegate certain tasks to certain models or harnesses.
aroman··on Fable 5 Is Back
This makes me think they really are quite capacity constrained at the moment.

I had assumed they were primarily limiting it to entice people to upgrade, but I feel like these limits are so low and so temporary (especially over July 4th weekend in the US) that people will barely get a chance to get "used to it" and then think: "man, I can't live without this, I'll pay for API pricing".

aroman··on Apple raises prices of MacBooks, iPads
I'm not sure I follow - 614 GB/sec is pretty squarely in dGPU territory (~5070 level). External GPUs can definitely exceed that on the very high end, but it seems pretty competitive, no?
aroman··on Apple raises prices of MacBooks, iPads
For sure, on paper - I'm curious, do you actually notice that difference in your day-to-day? I struggle to think of times in my usage of my computer where I think "this feels slow", but maybe I'm blind to it.
Page 1 of 27Next →