Seems like we've reached the event horizon of whether AI advances are worth paying attention to.
Not for me, Fable refuses to debug Linux kernel bugs. Unless you say who you're speaking for, it sounds like you're just shilling for Anthropic.
I use Sol and Grok 4.5 as my inline debuggers/reviewers, and both do well, and are decent at token save. DeepSeek V4 Flash 0731 found some interesting bugs when I tried it a few days ago, and I'm curious to see if that also joins the code-review line up
I recommend opencode or something akin to it to play with models. Any big model updates or hot new ones will naturally run across your desk that way
Spark is actually the interesting one imo. It's significantly better, also significantly faster. If you are ok with letting Meta soak up your data (which DS does too) it's also the same price.
A chinese model being in the same ballpark of capability at half the price sounds believable to me.
DeepSeek just spend almost 2 hours trying to figure out why terrain textures were not working. It tried everything over and over again, it even had reference code for meshes on how to setup the rendering with materials, and it could just not do it.
I finally gave up and gave it to GPT-5.6 Luna instead, and figure out in a single prompt after 20 seconds, that the terrain mesh was being initialized with None in the material slot.
Other tasks it has managed to figure out at least, but it is significantly slower than GPT-5.6 Luna and it requires a lot more iterations.
(Both were set to high reasoning)
Reading the DS reasoning is wild, it's constantly going in circles. The most minor lack of clarity in your prompt and it will spend ages going back and forth on what you meant. It reasons 5x longer than the preview which makes it really slow now as well. We did a lot of work to nudge it to be decisive and improve our evaluation setup to there's more clarity, and it helped but only marginally.
Ours is a full-stack app one shot test so it includes backend, frontend, design, and QA/testing. It's graded by Opus xhigh and Sol xhigh and the grades are averaged.
DS4 preview would finish in 20 minutes flat on high reasoning and grades 6/10. Luna high gets 9/10 in about 30 minutes. DS4-final is crazy - at high thinking it's taking over an hour and getting ~8 but only had one successful run as I got tired of waiting so long after many early abort/retries trying to debug why thinking was so long. The lowest thinking still takes over 45 minutes, and with thinking off it actually is finally closer to preview in time but actually get's a much more varying result anywhere from incomplete to 6 it seems.
Costs per run DS4 is best but not actually by a lot as it's spending 10x the tokens with all the reasoning and mistakes. It's a very brute force model and I really preferred preview in many ways for how predictably fast it was.
Side note, Spark 1.2 is a nice model for this test, best in frontend design and fastest to get results together, though not nearly as efficient as Luna. Grok scores similarly to Spark but at like $50/run vs the contributor Spark costing $1.50.
Edit: was curious to see and seems DeepSWE agrees at least: https://www.together.ai/blog/deepseek-v4-flash-0731-vs-gpt-5...
Edit 2: btw it tests a team of agents working together in a special harness and stack. So 20 minutes is for 8 agents essentially. That said everything was built around DS as it was the cheapest to iterate against so even with that advantage the new one struggles.
it's still $3/$15 for all providers on openrouter
because of some Kimi license
Uptime looks crap, though.
So we won't see any price decrease unless Kimi changes the license of K3
My read is, OpenAI is neither able to claw b2b money (away from Ant) nor are they able to stave off open weights on the other. In short, they're struggling to hold onto their distant #2 position in the coding market, and these pricing changes reflect a (desperate) change in strategy.
the private endpoint costs 10x (azure).
private endpoints for deepseek (lots of providers) also cost about 10x more.
but 10x more for deepseek is $0.028 cached input, and 10x more for luna is $0.10.