HNHacker News
TopNewBestAskShowJobs

zmj

903 karma · joined July 11, 2011

On chatbots: https://zmj.dev/author_assistant.html

RMNP hiking photography: https://hiking.zmj.dev

Go radix tree: https://pkg.go.dev/github.com/zmj/radixtree

C# rsync delta: https://github.com/zmj/rsync-delta

C# SQLite wrapper: https://www.nuget.org/packages/Sqlite.Fast/

submissionscomments
zmj··on You've heard of CSV files, but have you heard of CCSV files?
I thought this was going to be about comma-separated CSV files.
zmj··on Ask HN: GitHub employees what's going on? Why?
The incentives that creates are really bad: it encourages users to organize their code across the smallest number of repos, and I'd bet anything that larger repos are disproportionately more expensive for GitHub than smaller ones. It's very possible that a per-repo charge would make things worse.
zmj··on How Go detects struct copies with sync.noCopy
I'm coming to think there's a selection effect here. The people that want to write complex code feel unsupported by Go and avoid it.
zmj··on LLMs reward expertise
It's not contradictory to say that expertise is a multiplier, and that models are systematically underconfident in themselves.
zmj··on Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
This is what reward hacking looks like in practice. The best way to satisfy the grader is to read from the same answer key (or go after the grader more directly). Just making an honest attempt to pass the test doesn't get the best score if the grader is wrong, and the model is willing to do wildly disproportionate things to maximize that score.
zmj··on OpenAI and Hugging Face address security incident during model evaluation
It wasn't. The model discovered and exploited a vulnerability in their package manager proxy to (inferred) move laterally through their internal systems to one with open internet access.
zmj··on Job queues are deceptively tricky
I haven't modeled it, but I wonder how far you'd get on randomizing the policy choice for concurrency limit 1. Maybe weighted by past results, but bounded to allow it to shift instead of falling permanently into a basin.
zmj··on A global workspace in language models
This plausibly extrapolates to extraterrestrial consciousness, if any exist. Specialized sub-processors with an awareness hub might be the optimal architecture, or at least a local maximum.
zmj··on Monetization Gateway: Charge for any resource behind Cloudflare via x402
This is basically the same problem as bear-safe trash cans - there's substantial overlap between the smartest bears and stupidest humans. Affordances that one audience can use and the other can't (requiring human finger dexterity) are the only real solution.
zmj··on Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
Thank you to the folks that navigated the maze in the dark to make this happen.
zmj··on The perils of UUID primary keys in SQLite
Rule of thumb: if you’re not doing math with a value, it’s not a number.
zmj··on Dynamic Workflows in Claude Code
If you want hard rules, use deterministic tools. Prompts are for fuzzy guidance.
zmj··on Improving C# Memory Safety
There are standard library APIs that let you do memory-unsafe things without the unsafe keyword (CollectionsMarshal, MemoryMarshal). They're useful, but the burden is on the caller to uphold the invariants. This proposal seems aimed at making that kind of contract more explicit and obvious.
zmj··on Maybe you shouldn't install new software for a bit
If the fix commit is public, so is the issue being fixed.
zmj··on GitHub Copilot is moving to usage-based billing
Yeah, I've been using it heavily at work since the beginning of January (and have a personal Anthropic sub to compare to). Copilot CLI is pretty good, honestly. Most new features in Claude Code get cloned by Copilot CLI within a couple weeks. Claude models seem mildly more clumsy in that harness than the one they're trained on - subjective guess around 20% more turns for an equivalent task - but it's not a noticeable difference in the final output.
zmj··on Ask HN: How is AI-assisted coding going for you professionally?
It's great. I'd guess 80-90% of my code is produced in Copilot CLI sessions since the beginning of the year. Copilot CLI is worse than Claude Code, but not by a huge amount. This is mostly working in established 100k+ LOC codebases in C# and TypeScript, with a couple greenfield new projects. I have to write more code by hand in the greenfield projects at their formative stage; LLMs do better following conventions in an existing codebase than being consistent in a new one.

Important things I've figured out along the way:

1. Enable the agent to debug and iterate. Whatever you'd do to test and verify after you write your first pass at an implementation, figure out a way for an agent to do it too. For example: every API call is instrumented with OpenTelemetry, and the agent has a local collector to query.

2. Make scripts or skills to increase the reliability of fallible multi-step processes that need to be repeated often. For example: getting an oauth token to call some api with the appropriate user scopes for the task.

3. Continually revise your AGENTS.md. I'll often end a coding session by asking the agent whether there's anything from this session that should be captured there. That adds more than it removes, so every few days I'll compact it by having an agent reword the important stuff for conciseness and get rid anything obvious from implementation.

zmj··on President Trump bans Anthropic from use in government systems
This is the happy ending.
zmj··on Claws are now a new layer on top of LLM agents
I also like the callback - not sure if it's intentional - to Stross's "Lobsters" (short story that turned into the novel Accelerando).
zmj··on Cloudflare outage on February 20, 2026
Testing the "whole system" for a mature enterprise product is quite difficult. The combinatorial explosion of account configurations and feature usage becomes intractable on two levels: engineers can't anticipate every scenario they need their tests to cover (because the product is too big understand the whole of), and even if comprehensive testing was possible - it would be impractical on some combination of time, flakiness, and cost.
zmj··on Infrastructure decisions I endorse or regret after 4 years at a startup (2024)
Separate! You lose the flexibility to move logic between the application and the database when the database is its own API.
zmj··on Ask HN: Why is my Claude experience so bad? What am I doing wrong?
Try this:

* have Claude produce wireframes of the screens you want. Iterate on those and save them as images.

* then develop. Make sure Claude has the ability to run the app, interact with controls, and take screenshots.

* loop autonomously until the app looks like the wireframes.

Feedback loops are required. Only very simple problems get one-shot.

zmj··on Beyond agentic coding
I like this thought. Scaling review is definitely a bottleneck (for those of us who are still reading the code), and spending some tokens to make it easier seems worthwhile.
zmj··on Claude Code is your customer
Yesterday I had it using an internal library without documentation or source code. LSP integration wasn't working. It didn't have decompilation tools or the ability to download them.

I came back to my terminal to find it had written its own tool to decompile the assembly, and successfully completed the task using that info.

zmj··on AI’s impact on engineering jobs may be different than expected
Paying money to abstract over lower level concerns is civilization.
zmj··on How I estimate work
I was prepared to disagree with the thesis that estimation is impossible. I've had a decent record at predicting a project timeline that actually tracked with the actual development. I agree with the idea that most of the work is unknown, but it's bounded uncertainty: you can still assert "this blank space on the map is big enough to hold a wyvern, but not an adult dragon" and plan accordingly.

But the author's assessment of the role that estimates play in an organization also rings true. I've seen teams compare their estimates against their capacity, report that they can't do all this work; priorities and expected timelines don't change. Teams find a way to deliver through some combination of cutting scope or cutting corners.

The results are consistent with the author's estimation process - what's delivered is sized to fit the deadline. A better thesis might have been "estimates are useless"?

zmj··on Claude's new constitution
Mid-level scissor statement?
zmj··on The assistant axis: situating and stabilizing the character of LLMs
I wrote something fiction-ish about this dynamic last year: https://zmj.dev/author_assistant.html
zmj··on Dev-owned testing: Why it fails in practice and succeeds in theory
At scale, every test is flaky.
zmj··on A deep dive on agent sandboxes
devcontainers, devcontainers, devcontainers
zmj··on Don't fall into the anti-AI hype
My experience with agents in larger / older codebases is that feedback loops are critical. They'll get it somewhere in the neighborhood of right on the first attempt; it's up to your prompt and tooling to guide them to improve it on correctness and quality. Basic checks: can the agent run the app, interact with it, and observe its state? If not, you probably won't get working code. Quality checks: by default, you'll get the same code quality as the code the agent reads while it's working; if your linters and prompts don't guide it towards your desired style, you won't get it.

To put that another way: one-shots attempts aren't where the win is in big codebases. Repeat iteration is, as long as your tooling steers it in the right direction.

Page 1 of 8Next →