172 karma · joined July 11, 2023
I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal.
Luna is similar, and also 8x cheaper. Source: artificialanalysis
The only benefit I can see is the speed, that looks to be outstanding, probably thanks to their TPUs.
Inference will be always more expensive than db operations or copying, sure. But how much more expensive is the question.
No. There is economic opportunity cost (borrowing), energy cost, infra cost, depreciation / risk of failure with each unit of work, bandwidth, maintenance, and lots more. Small, but not zero, and often overlooked - especially the opportunity cost.
It’s so cheap that companies choose to spend more on AI inference (more reasoning, more capabilities, longer context), not less - see Jevons paradox.
Of course, and so does everything in the software world. The point is getting the cost so low that it’s basically free. The new DS V4 Flash or the smaller Qwen3.6 models are still really expensive compared to what we were used to in the economics of software, but it’s not unreasonable to expect these costs to continue falling down.
Rough chatgpt estimate says 3-5 orders of magnitude of difference compared to a typical user interaction with a SPA (db/cache lookup, CDN…)
- Haiku: 30 points
- Luna Medium/High/Xhigh/Max: 38/46/49/51 points
That's a massive difference:
- 30 points is Gemma 4 31B territory
- 50 points is GLM-5.2 (744B) territory.
But this whole post seems a bit fishy to me. Brand new account, and it starts with "Title:" and "Post:", the whole thing being obviously entirely AI generated, and a few other signs.
Just a thought, have you tried any way to triage these reported issues via LLMs, or constantly running an LLM to check the codebase for gaping security holes? Would that be in any way useful?
Anyway, thanks for your work on opencode and good luck.
https://openrouter.ai/announcements/response-healing-reduce-...