He did clarify that it was with fast mode. Without fast mode it'd "only" be $300k in raw API cost, or ~60 $200 Codex subscriptions.
He did clarify that it was with fast mode. Without fast mode it'd "only" be $300k in raw API cost, or ~60 $200 Codex subscriptions.
Business: Amazing, that’s great what did you do?
I ran 50 instances and had them all fix the same bugs at the same time and then analyzed the results of all 50 runs to have AI score each of the attempts, then sort them, then compare them to each other in a round robin tournament style double elimination to ensure I got the best result. Then I had AI convert this into a skill, and then ran all 50 attempts again and repeated the process to ensure that I had the absolute best result. It was amazing and I used 1.3 billion tokens!
Business: That is amazing! What did you fix?
A spelling mistake on the About page.
Wish me luck on a raise!
Eventually Codex's subscription subsidization will diminish to near-zero, like the rest of the providers.
It's extremely important that people understand how expensive these models currently are. Even $300k in raw API costs is alarming for the output.
Because it does not say “equivalent of”, it literally says he spent money that he did not spend
So yeah its misleading but in the other direction.
He has agents write shitty code for features other agents think other people want, then has it reviewed by other agents in hopes of catching bugs that the first agent put there, then has some more agents try to find security bugs in the now double-agented code to make it triple-agented and at the end of the day, he spent a shitton of tokens, probably emitted enough carbon to heat our planet by another degree, and has a feature nobody really asked for that might or might not work.
He then has the sense of humor to call this grotesque process "incredibly lean".
What's the point in all of this? What problems is this solving? Who's benefiting?
> What's the point in all of this? What problems is this solving? Who's benefiting?
The economy doesn't work like how you think it does. Its not central planning. All the usages aren't detailed in a specification, submitted for approval to 100 agencies and then allowed to be used.
It shows lack of intellectual curiosity to not engage deeply with obviously profound technology and what the implications are. I find this exercise helpful.
Peter is predicting how LLMs will be used in the future when the prices go down. And they will definitely go down. I think his predictions are correct and we will definitely have something similar to OpenClaw.
I didn't know that studying photocopiers is suddenly linked to "intellectual curiosity". Being a photocopier maintenance guy was always considered boring.
What you put on top of the machine was intellectually interesting.
I'm aware. That is in fact my central critique. The way it works is incredibly wasteful of our limited resources, as illustrated by this guy burning through fuel during a time of crisis for no perceptible gain.
> It shows lack of intellectual curiosity to not engage deeply with obviously profound technology and what the implications are.
The "obviously profound" is an assertion without proof.
The rest I agree with, we should engage with the implications of burning through energy to build features that bots think humans want, but nobody actually asked for, all while climate scientists are telling us we're heading for the apocalypse. It is intellectually incurious to just ignore the questions of why and at what cost, maybe even dangerously so.
You should try playing the game “workers and resources”; it’s a simcity like game, but based in the Soviet system of central planning, not capitalism. It will make you loathe the inefficiencies in central planning.
The appropriate comparison is command vs market. Capitalism is efficient in utilising the characteristics of humans to bring about expansion of markets.
like one bot finding similar issues and PRs, the another bot closing issues for "lack of activity", meanwhile people are reacting and pleading to speak to a real human?
Congrats builders of the future, you've turned software development into automated voice systems.
“He has /people/ write shitty code for features other /people/ think other people want, then has it reviewed by other /people/ in hopes of catching bugs that the first /people/ put there, then has some more /people/ try to find security bugs in the now /double-peopled/ code to make it /triple-peopled/ and at the end of the day, he spent a shitton of /money, the people/ probably emitted enough carbon to heat our planet by another degree, and has a feature nobody really asked for that might or might not work.”
Honestly sounds like a normal tech company to me. Just with much dumber “people” who are getting exponentially smarter, eventually never die, eventually never forget.
You have to skate to where the puck is going, not where it is.
They haven't gotten any smarter yet, let alone exponentially smarter. They are still the same dumb parrots that they were in the beginning.
The morality issues about consumption climate impacts are not his alone, and are not unique by itself to his endeavor. Every company with an enterprise LLM agreement has a share, for instance.
Firstly, who TF would use that crap in the first place at all? Yeah, he did some crap he got paid for. So did the people who created the addictive algorithms for social and media or creators of the brainrot videos that infest kids' minds. Should we applaud them too?
The reality is if the thing can’t survive financially without being subsidised in the long run - it deserves to die.
There is no trend showing that these expensive things exist in the long run in this manner. - it’s pure speculation and for many: hopes and dreams.
Who knew it was that simple..?
It's a very simple question, the subthread you created based on reducing everything to "he did a thing" and calling the comment you didn't interact with at all "agonizing".
Why not rather leave it at "they wrote a comment"? What is so hard to understand about that, to use your words?
The money going to the American model companies is not going to their hosting costs.
What I mean by this:
1. Intern, analyst, junior, or offshore level coding is cheaper when done by the machine.
// Side note: There is good reason the industry invests in suboptimal output from this set which moves to the "cost" column when using an LLM, but nobody's accounting for that.
2. For the interns, analysts, junior, or offshoring to do the right thing costs a multiple of the coding effort: the PdM/PjM stuff of course, but also the Stakeholder, Product Owner, Architect, Principal Engineer, QA, and SRE stuff.
3. If you are not a principal or staff engineer level engineer, you are likely unqualified to catch and fix the errors LLMs make across engineering, much less these other PDLC (product development lifecycle, which includes SDLC and SRE) loop.
4. For LLM output to be useful, your 'harness' has to incorporate all of that as well, which because it's so much harder than transliterating spec-to-code, balloons tokens exponentially.
5. Today it is faster, more efficient, and costs less, to work with LLMs "XP" (eXtreme Programming) style, pairing with the LLM actively co-creating and co-reviewing, steering for more effective turns.
So, your options are:
- ship garbage while costing less than a median first world SWE
- pair with the LLM actively for the benefits of XP
- add enough harness and steering the LLM costs more than SWEs, and still needs a human loop “move fast and break things to find out what's broken” style
I would expect that within a couple years, these other disciplines can be baked in enough the machine costs less for everything but surprises.
They already are. I’m successfully using frameworks like bmad to deliver complex apps at that level. My job is to manager the see, as, ux, sre processes and catch errors.
I spend more time refinding prd , epics and stories than I do elbows deep in code.
If I don’t like the output of a story I nuke it change the story and have the flanker try again. I’m using the open source glm, kimi, deepseek models. I expect the full pipeline to be good enough by the end of the year.
And do you enjoy this more than writing code? I used to look forward to writing code, solving these little optimization puzzles, learning, and staying sharp. Working with agents is dreadful in comparison. They lie, rarely learn, and I feel like a proctor.
Sure, you sometimes get to see something amazing, but usually I am just very annoyed by their performance and ever-changing but never-ending billing issues. First, with Claude Code, now with Codex, which was fine for a minute, but now I am out of tokens for the majority of time. (I don't have the income for those Pro INTx plans.)
Now, I'm master of about a thousand lines of pricing code plus documentation and research which actually matters. The AI can handle the rest as a very skilled junior with a TBI.
You absolutely are a proctor, or senior manager. The AI is the smartest most well read junior you will ever meet, but don't go out of its happy path.
As you go out of the commonly read happy path for CRUD apps, you'll have to get more and more involved. I wouldn't write a new kernel design with AI right now, I might write a Linux kernel driver with it though.
ya'll cant have it both ways; either it's really worth the cost or it's a bunch of token burn with no smoke.
They literally are. (If by "all this" you mean the subscription future bait-and-switch plans.)
Lets say I was at the casino and was spending a lot on casino chips but I also happen to work at the casino. I'm not really losing money whether if I win / lose since I'm using the houses money and there's little risk involved on every dice roll or press of the button. The risk is far higher if I don't have that level of access and continue to spend the same amount of money on lots of tokens (or casino chips, spins or button presses.)
The same is true here with these agents. Some companies will realize that they can no longer afford to spend millions a month on tokens or even startups spending $5k - $6k per person per month on tokens.
I can only see local efficient models making sense on recovering from this unnecessary spending or even light gambling on tokens.