A verification loop 4x'd DeepSeek's intelligence, matching Opus at 1/7 the cost
ironbee.medium.com
ironbee.medium.com
The skills can be found here:
https://github.com/ironbee-ai/ironbee-devtools-skills/tree/m...
Question that was left unanswered:
How much longer does a 4x DeepSkeep loop took compared to Opus?
And if this is worth a look, I’m sure I’ll hear about it again from someone who wrote it better.
Posts like this break the social contract, and honestly show a lack of both care and consideration for communicating the subject matter AND a flagrant disrespect for anyone who would consume this.
I’m sure there are some good ideas here, but it’s hard to sift through the lack of applied ai and lack of communication skills to get to them
I hope you can take your own advice?
Take care!
I don't need to eat entire bowl of the burnt soup with flys floating in it to know its a burnt soup with insects in it.
Why do you expect me to spend even more time to laboriously show you examples of the slop? The other comment already did, I only engage with the actual work, I wasted too much time on it already.
Better luck next time. Oh, and a free advice, when someone rejects something you posted, a thing, immaterial, and you recoil with personal attacks at anyone who didn't like it, this is one of the ways you show it's not worth investing more effort into discussing with you or your agents.
Instead of
> Is the loop just more attempts? One objection deserves an answer up front. A verification loop spends extra inference per task, so is the lift just a bigger compute budget in disguise? Partly it has to be. The loop does more work. But the retries the benchmark grants are blind: the model sees a failure signal and guesses again. The loop’s iterations are guided by evidence from the running application, which is a different kind of attempt, not just another one. Whether guided iteration beats an equal budget of blind retries at matched cost is exactly the ablation this framing demands, and it is planned for a future post: DeepSeek alone with a larger retry budget, against DeepSeek with the loop, dollar for dollar. Until that runs, read the results below with this open question in mind.
It could have been
> These loops are not just retries. Each iteration provides the model with evidence from the previous run. Some of the uplift may come from the extra tokens, so a follow-up post will compare the guided loop against cost-matched blind retries.
We can argue over exact wording, but the original is far too long.
Or the point about "measuring cost honestly". It's not clear why you wouldn't be using the published rates and do the basic multiplication yourself. There's nothing subtle about this, and it doesn't need to a whole paragraph.
> Or the point about "measuring cost honestly". It's not clear why you wouldn't be using the published rates and do the basic multiplication yourself. There's nothing subtle about this, and it doesn't need to a whole paragraph.
Because, while Opus provides cost in the API response, DeepSeek doesn't, but just token usage. So, cost is calculated based on used tokens for DeepSeek. We have added this section, because draft version of the post got feedback on this.
If someone vomits the slop that must be fed to a robot to make sense of it, even if there's something of value.
It's just so rude, like a professor entering the class, emptying a carton box full of papers on the floor, flipping a bird and leaving without a word.
It's possible there's something good in it.
I said I didn't read (so much for the big gitcha) because if its not designed for human consumption that means it's not worth human time.
I bet someone actually respecting their audience will present findings in a way that's.... palatable, for the lack of the better word.
(I spent some time commenting to add a voice of disagreement for enshittification of the publishing part of the internet. I'm entitled to it as much as you feel entitled to posting your comments)
Haters like you, who skim other people's comments and pass them off as their own thoughts, are everywhere. I don't have time for it. Try you chance with other posts.
I didn't say I skimmed your comments, no need to lie if you don't have a good response.
You're too enraged people expect some baseline quality from whatever you post. Why is that so much?
And also, I'm not saying you said you skimmed my comments. It doesn't need to be said, it is obvious that you wrote those comments without opening the link by just reading other's comments. That's it!
That's the neat part. You don't... need to do this yourself nowadays. AI agent is perfectly capable to figure out any new open source thing that pops up daily now. You can just give it the link to the repo and say "Set this up for me. I wanna play with it."