3,427 karma · joined March 5, 2014
Unless you're a real contrast freak, I think most people would be slightly happier with the color than the B&W; but it's definitely not worth paying extra for.
Those people are doomed. Which is fine in a lot of ways, but will be devastating for their economic prospects.
Also, though, I think some of the weird poverty-culture stuff of the Netherlands is just a culture thing. Netherlanders are absolutely rich enough to buy paper towel if they wanted to, it's just a cultural thing that they don't. And air conditioning is a whole cultural hot button flashpoint issue for reasons that have nothing to do with money.
They also have an update -- linked from the original study! -- explaining that it's out of date and no longer reliable, and explaining why they had to cancel a follow-up study because it was understating productivity gains (but also was showing wins for the people who carried over from their previous study): https://metr.org/blog/2026-02-24-uplift-update/
The authors of this paper decided to ignore all of METR's follow-up data and discussion, and to report only the ancient number from early 2025 (a time when Windsurf was state of the art). And then, rather than apologizing for it, and caveating it as a number not to be taken seriously, they described it as a study done "recently."
That's either shockingly dishonest or incredibly out-of-touch.
That meeting that you spent an hour in to understand the requirements? You don't need that meeting if you're not writing the code. That sync up with the QA engineer you did to hand it off to them? Don't need that meeting if you're not writing the code. That half hour you spent installing vim extensions? Don't need 'em if you don't open vim anymore.
There are engineers whose jobs go well beyond coding, of course. Staff engineers and principal engineers have had their jobs radically change because of AI, but not because it's writing all their code.
But there are also a lot of engineers -- your standard mid-level engineer, or even senior engineers at a lot of orgs with title inflation -- whose job is almost entirely about delivering code, and who spend all day either writing code or engaging in scaffolding around code-writing activities. Let's not pretend that automating away that code writing is a 15% boost.
California's current net metering system, which is more reflective of the true benefits and costs of solar to the power system, has been incentivizing people to add storage to keep their solar rather than get paid big money to feed back midday low-value solar, which is a big win.
If you subscribed to Netflix back then, you probably subscribed to both the streaming and the disc-mailing plan, because there were enough things you couldn't get on streaming, so getting the discs was important to fill those gaps.
Yeah, it's testing a different thing than what the benchmark claims to test, but it's also accidentally testing something more real-world applicable than a clean benchmark would be, so hey.
(EDIT: That is, if the agent is allowed to see the failed tests and iterate. If not, then yeah, that's just a problem. And either way, the ones with tests that just encode a particular solution's implementation details, thereby demanding that your solution have some rando internal details, are junkier. That's not a situation you'd run into in reality.)
(Not universally, but in many cases.)
But it sounds like you're really asking about the state of the world today. If so, I don't think that ideal state is like your friend's company (or at least, as it appeared to be to you). It might be possible that you can make that "dark factory" pattern work (StrongDM seems to be doing it), but it would require infrastructure and discipline that I doubt they're mustering. Think about how CD didn't involve taking a sloppy build process with no testing or observability and just going straight to prod -- it required building up a lot of infra and discipline first.
But on the other hand, I don't think the ideal present involves artisan hand-crafting code either. I haven't written a line of code by hand in enough months that it would genuinely feel weird if I were to try to program that way despite decades of having done just that. That era's done with, and moderate normie practices right now today are more about supervising and guiding agents than about chiseling code into clay tablets.
(I think they are being irrational, and that the mental model they have of AI costs -- "how much are we spending on tooling for this developer?" -- is going to shift over time to something more sensible, but those kinds of short-sighted companies are the ones that are having cost panics.)
In the past, a team of five mid devs and one good one would be fine, because that good one would ride herd on the mid ones. But now those mid ones are slamming out robot code that they're incapable of meaningfully reviewing (because it's better than they are already), and they're just overwhelming the good dev's capacity.
The solution, of course, is to fire them all -- they're worthless now -- but this is not going to happen quickly, and it's probably for the best that it doesn't.
(For as long as that's true, "software developer" is still a job. It's not clear for how long it will be true.)
But with the agent, you know that the change will be relatively quick and easy, so the bar to tell it to shift approaches is much, much lower.
(There are workplaces where that's the norm, I know -- it tends to be a thing with smaller teams with codebases that everyone understands fully, and much less a thing with larger teams where different people have areas of the code they understand more than others.)
With AI code, though, it's _your code_ and you can't give it a lgtm, you actually need to dig at it until you do fully understand it, fully agree with it, and could justify it to a hostile reviewer. It's a different level of rigor.
Not all engineers apply that rigor, though, which becomes a problem.
But trying it out... alas, no. Simple factual questions where ChatGPT would go do a quick search and get the facts and report them back to me, get a "Great question! [totally invented bullshit]" from Claude, even with this new model and thinking set to high. I have to explicitly tell it to search to get it to look up basic facts, rather than it recognizing that it needs to do that, like GPT does.
If people would do even a little bit of math, they'd see that Microsoft can't possibly be paying more for AI than for developers: They have about 80K employees in product development roles. Senior developers probably cost them $400K all-in.
Do they have a $32 billion Claude bill? I suspect they do not.