357 karma · joined September 29, 2025
Bonus forth point: why is this critical to solve for claude code, but not for all the other harnesses which have all converged on AGENTS.md for this purpose?
2. How exactly do our jobs depend on a thing which has been around for far less time?
1. The resulting code was of pretty low quality. Others that bothered to put it through Miri and the like found many soundness issues, but my personal favorite example is this example which is trivially and locally (meaning that someone who has the most basic understanding of unsafe in rust can see it's obviously wrong just by looking at the specific function) incorrect example [0]. This particular example was removed in an apparently unrelated refactor after spending well over a month in the code base without any of the bun maintainers or their agents detecting it, and a quick grep found hundreds of potential similar issues (although many of those are false positives).
2. More generally, it's not clear to me that there was any technical benefit to the rewrite in the first place. The stated reason was for memory safety, but replacing Zig with unsafe rust doesn't actually get you memory safety, and removing the unsafe blocks often requires more extensive refactors to fit within rust's model.
> If you can reduce a problem to a clearly verifiable end state, provide the necessary context, and equip a model with the necessary tools it can usually get to a good solution.
As others have pointed out (and you acknowledge), "reducing a problem to a clearly verifiable end state" is just "programming". What you don't seem to understand is that actually doing that is made harder by using AI, not easier. A sufficiently detailed spec is called "code" [1], the question is what language/notation is best to write it in. The answer is almost never "whatever is closest to what the computer actually executes", as assemblers and later compilers and interpreters demonstrated. But it also isn't several of the things that AI proponents have suggested to replace the latter with.
Take natural language, for example. As Dijkstra pointed out, we've been through this already with math. It used to be that all math was expressed in a way closer to what we'd now call "word problems", but this turned out to be bad. The specialized language of e.g. algebra isn't something mathematicians use to gate-keep, it's way easier to reason in the domain that way than it is in English (or other natural languages). The same is true for programming, once you actually specify what you want to do with enough rigor. It's generally easier to read and reason about code than to do so with natural language specifications.
Another proposal is to use tests and similar automatic verification to specify the program. I suspect that anyone with much experience can already tell whether it's preferable to specify a program through code or through tests, but thankfully we have empirical evidence on this for anyone who has any doubts in the form of e.g. sqlite. Sqlite is probably one of if not the closest any piece of software comes to being fully specified by it's tests. To do that takes almost 600 times as much test code as there is "regular" code. Dr. Hipp even personally weighed in on the implications this has on AI recently [3] . The reason to do testing is that it provides a second independent check for correctness, if you're using it as the *only* check that advantage disappears.
[0] https://github.com/oven-sh/bun/blob/fc865b398e51de8a95ddde4b...
[1] https://haskellforall.com/2026/03/a-sufficiently-detailed-sp...
[2] https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...
LLMs are incredibly impressive. It just doesn't follow from that that they're capable of all the things proponents claim they are.
> I just can't believe that LLM haters really like computers or technology
In a lot of ways the reverse is true. The LLM coding ecosystem seems designed by people who are unaware or actively despise the fact that they have access to a computer, and almost the entire point of LLMs is to make interacting with computers less like interacting with computers and more like interacting with people.
What about the people who do have a lot of LLM experience and agree with OP?
> That thing cannot do the stuff you say it can. And there is no way I will ever try it. But I know for sure I am right.
If someone claims to be able to fly by strapping bird wing shaped pieces of plywood with feathers glued on to their arms, do you need to personally jump off a tower with them to say they don't work, or can you look at the results of others attempting it and draw conclusions based on that? LLM proponents are making claims about their capabilities which can relatively easily be checked without using LLMs yourself. To pick an example where LLM proponents are correct, anyone who says LLMs can't generate syntactically valid code can be proven wrong fairly easily by producing an example of syntactically valid, LLM generated code.
Further, if you read the rest of the paragraph you responded to, it's clear that this claim is about the state of the industry as a whole. "Software has not improved in quality, got faster, become cheaper to produce (when you exclude the mountain of poor-quality demoware that no reputable organisation would touch with a barge pole), or become more capable." Whether this is true or not is something that can be evaluated without ever having prompted yourself, or even arguably without being a developer at all.
If you use google (sans AI), you're putting some trust in their page ranking algorithm. If you use it with AI, you're trusting the same algorithm (since that's how the model gets it's sources), but then you're trusting the model to evaluate the sources for credibility and extract the information you actually want.
[0] https://www.techradar.com/computing/cyber-security/facebooks...
[1] https://www.tomshardware.com/tech-industry/artificial-intell...
[0] https://businesschief.com/news/why-are-executives-using-ai-m...
The creator of HTMX would disagree with you (as the person who described and coined the term REST)
https://htmx.org/essays/how-did-rest-come-to-mean-the-opposi...
Another point of comparison is density of unsafe: the number of unsafe blocks per line of code and/or file. By this metric, Deno has a bit over half the unsafe (because the bun rewrite is significantly more lines of code).
First off, you seem to be under the impression I'm a rust hater. Noting could be further from the truth. Rust is easily my favorite language at this point, I reach for it for basically everything (except quick scripts). While I do like a lot of zig's philosophy, I think at the end of the day the empirical evidence is overwhelming that manual memory management isn't sufficient.
> What's your point, that if we can't do everything perfectly in one step we can't do it at all?
My point is exactly what I initially said: you typically aren't much closer to a (mostly) safe rust codebase if you've done a line by line port to (partially unsafe) rust than you were to start with. Getting to safe rust is very likely to require substantial refactors either way. This doesn't mean you shouldn't do it (on it's own), but it does mean that the bun team's strategy/assumptions are more questionable than they appear to realize.
But second, you're right, this is an easy thing to search for (especially with LLMs). And yet, the example I linked has been there since may, surviving multiple rounds of review. From this, we can draw a few conclusions: 1) claude (including apparently mythos/fable) fundamentally doesn't "understand" how unsafe rust works, and 2) no one on the bun team is aware that they need to tell it to fix this (if they're even aware that this isn't allowed in the first place).
The reason such trivial defects in a codebase are a red flag isn't so much those specific issues themselves, it's what they reveal about the authors. If you find a lot of obvious defects in some code, it's very likely that there are also many more subtle harder to detect issues as well.
[0] https://github.com/oven-sh/bun/blob/fc865b398e51de8a95ddde4b...
Your hypothetical developer wouldn't be using notepad because they're unaware of other editors, they'd be using it because they evaluated other editors and concluded that, for whatever reason, they would be worse for them. I'd be fascinated to hear why they came to that conclusion, but I'm not going to tell them they're wrong if they're performing acceptably, aren't constantly breaking CI because the linter rejects their code, etc. Everyone is different, and I'm not narcissistic enough to think the fact that I would be way less productive without my modal editor, LSP, linter, terminal multiplexer, etc. justifies forcing everyone else has to adopt my exact setup.