The joy of watching a dumb AI-ism be sharply corrected by code you wrote months ago is hard to explain.
2,723 karma · joined December 1, 2014
The joy of watching a dumb AI-ism be sharply corrected by code you wrote months ago is hard to explain.
Thanks to everyone who worked so hard to get here.
My plan downgrade kicks in next month.
Thanks Opus 5 for helping me kick the habit.
Tim (author) if you're there: it'd be amazing to see the split of which agents people are using, if you have that data.
It's like making a beautiful chocolate praline and shipping it in a shoe.
Running a test, then running it a second time, then asserting something else.. is just weird.
4F is not enough to rise above a noise floor; 4C is getting close.
If 2026's Anthropic did an announcement like that, it'd be so many words it'd crash the browser.
"we seem to keep shooting ourselves in the foot, and the solution we have invented is bulletproof shoes"
But your first 3 paragraphs serve no one but you.
My suggestion for your future, if you'll consider it, is to write those first 3 paragraphs (they feel good to write!), but next, take a deep breath and delete them.
I agree with your point by the way - everything has pros and cons - and there're tipping points. "Too much of a good thing" is true of so much in life.
Better call your AI company for a refund. That my friend was an hallucination and you've been promised no such thing.
Hm. Two orthogonal properties! This sounds like a 2x2 matrix!
Let's swap hard/easy around & explore the 4 possibilities...
There are domains with sharp delineation and hard verification; they are not at risk until AI gets much better. Humans operate in these domains by applying tremendous deep thought and subjective judgment - our superpower.
Domains with soft delineation and easy verification are most at risk: "it's a picture of a cat" remains true through a wide range of perturbations - eg. skewing the image or moving it across a pixel or correcting its white balance or even changing the cat. AI music? Lots of domains already solved by AI here but they're also not that meaty.
My prediction is the next interesting stuff will happen where verification is hard but there's no sharp delineation. It's the world of "I'll know it when I see it". Good customer service?
> "Please add in this repo, a pre-commit hook in the form of a script that prints all filenames and line numbers on which the hyphenated "open-source" string exists, and if any instances are found, will exit nonzero. Then please install it in `.git/hooks` so commits are blocked until me or an agent removes all the instances, so we never again commit the string "open-source"."
1 minute after hitting enter you've created yourself a guarantee your repo will never again see the string "open-source".
No amount of tokens can come close to my hourly rate.
Ground your LLM. Tests, documentation, give it many ways to run the thing its reasoning about. It needs to be able to test its hypotheses on its own.
Take yourself out of that loop so you only find out once it's sure.
Do you have evidence?
I agree it's incomplete what I wrote; you were right to call it out.
When you as an individual are getting tens of thousands of dollars of value from a $100 item and someone says there's a slightly-worse version for $1 but can't tell you if you'll find it slightly worse or noticeably worse.. you have to stop what you're doing to evaluate them, pay switching costs etc etc. so the cost difference is less than $99 and might actually be negative, and the difference in margin after paying switching costs is negligible.
So your rational response is to investigate to lower the stakes - but even performing the investigation eats into the time you're spending getting your tens of thousands of dollars of value.
Below a certain threshold prices are effectively all the same.
Companies pay API prices because part of the bundle is a trust anchor / liability shield: "we bought Anthropic's thing! We didn't risk it! We paid for ZDR!"
It's not an expensive car vs. any car.
It's car vs. no car.
It's as easy as ever to lose yourself or get lost if you so choose. Probably always will be.
API prices are paid by companies getting tens of millions of dollars of value.
In normal life money is the key constraint; buy this don't buy that etc. - whereas in VC funded companies the constraint is time. If you as a founder get funding and don't spend it fast enough you put yourself at serious risk of being replaced.
When enough of the world operates on that principle it creates a highly price insensitive market and that then can support a ton of ideas and experiments, some of which turn out to be really really good. It's a wild way to do innovation but it's been working well for decades.
You can get just as good information by asking its thoughts for and against some issue.
That doesn't force it to stop being sycophantic; in fact it actually exploits sycophancy to give you what you want.
RSS is amazing and being able to zip through your feed without waiting for pages to load is a game changer on top.
By way of analogy consider the relative impact on an ecosystem of one person fishing with a fishing line vs. a commercial fishing boat trawling the ocean. Of course, one person fishing is unlikely to have a huge impact on the ocean so it's generally permitted. Trawling (agentic coding) can be done in a way that's destructive to ecosystems but it can also be done sustainably!
So with the analogy in mind let's bring back the "sensible trawling" idea to agentic coding. What might it look like to solve the "huge PR bad" constraint in another way: by increasing our codebases' ability to absorb change, so what "a huge PR" is, becomes bigger?
Probably needs solves at many levels: assistance quickly comprehending the PR (AI driven walkthroughs, multiple media expected from the PR submitter not just text - eg. a screencast walkthrough of it), it requires rethinking how the code is read (better review tooling); it requires integrations with code-review automation tools (both you home-grown checklist and third-party tools) it requires rigorous testing (comprehensive automated e2e; test-driven; functional tests; etc); it requires putting the actual "in the loop" so post-release fast-follows can be expedited (eg. product signals and Sentry and metric anomalies are fed back in for quick follow up releases)
If you can be so much more responsive to the customer and market. Eg. you can unlaunch features just as easily as you launched them - and you can finally clean up all that tech debt. Better for the business better for the codebase and better for developer happiness.
Not all these strategies work for every situation - you can't do post-release in the loop if the shit needs to work first time! But the whole idea creates so much richness in applying human judgment and engineering solutions and it's all brand new because we never needed to deal with this much change before
Think of it as "releases in the loop".
If opportunities to rethink the stack to support MORE change excite you, congratulations! You're ready for the future that's coming. If you don't like this - get yourself into a job where you can say no a lot, or where shit needs to work first time, and you can be happy. Test-driven, strongly reviewed.. there's ways with agentic coding to also make super high quality stuff. But you can also shoot product from the hip more accurately and more often than ever before.
It won't be applied correctly everywhere - it's still heavily judgmental dependent and we're all fallible - but there'll be a much wider spectrum of options for how to build products. I think this is a really exciting future!