Now, if the debate is really about which option is more cost effective, then we could easily run an A/B test to find out. Though TBH my instinct is that that experiment is likely to cost more than the potential cost savings.
What I will say is that my own sense from experimenting around in a non-rigorous way is that the answer depends on how you use the tool. For actual vibecoding you should always go for the SOTA model because it will need less oversight. It’s also less likely to get stuck in a vicious loop that fruitlessly wastes tokens. But for a more hands-on approach where you move in small, carefully planned increments that you review and test in human-comprehensible chunks, smaller models may be preferable. SOTA ones don’t do that much better when working that way, and the slower inference adds a detrimental amount of friction to the work cycle.
I'd object to "easily". It's already hard to measure whether AI is generally worth the cost. Let alone compare models in such detailled ways. It's mostly handwaving and gut feeling Doesn't mean the conclusions are wrong, but biases are strong.
If this is a cost conscious company where I'm going to get a fairly limited amount of Sol, or a nearly unlimited amount of Luna, I'm probably choosing Luna.
Confidently wrong here. It's absolutely not across the board "more capable" than GPT 5 or Opus 4.1 or even Gemini 2.5 Pro. It's potentially better at certain specific tasks, mostly agentic coding implementation work. I.e. tool calling and usage of bash. It's worse at a large range of other tasks.
Sol medium has been a nice balance between intelligence and response time. Luna xhigh can achieve similar scores on the evals, but it takes noticeably longer. My impression is that the higher reasoning effort helps compensate for the lower base intelligence.
Cost is definitely a big factor, but latency and intelligence matter too. If I had the budget, I’d take Sol medium over Luna xhigh.
From using both on real scenarios, Sol is noticeably better at navigating around issues, exploring alternatives, and being creative when the obvious approach doesn’t work. That matters quite a bit when you’re investigating live alerts, where the path to the root cause isn’t always straightforward.
Sometimes I will use Fable or Sol for large features/projects, or research/exploration.
I would not be at all happy if I were forced to use Luna, though. I’d probably start looking to leave. I don’t want to work somewhere where I don’t have choice over my tools.
As long as there is no cost involved there is little reason to restrict such tooling, but expecting your employer pays non-trivial sums is likely to raise questions about whether that extra cost is worth it. In this case, is the delta from luna to sol worth the extra order of magnitude in cost. If it's $2 a day vs $20 a day then there should be little discussion, but at $50 a day vs $500 a day, the question seems relevant.
It is not reasonable for a company to pay a fully loaded engineer >$40k/mo and then ask them "why did you need $500?" on some odd days where they are doing difficult work, large refactors, experiments, etc.
Yes. Categorically. Anyone who tells you otherwise and that luna is “just as good” does not know what they are talking about.
Going from sol to luna is a downgrade.
It is not a question, it is a fact.
> Is sol actually worth the extra cost?
Is a question only you can answer, because it has no generic answer.
Right now, for me, being able to use sol is worth the cost, but using it all the time is not.
I’m sure going from using it to using luna feels rubbish; but there are realities about costs you have to face sooner or later.
Maybe like… give your team credits and make them pick the right tool for the job; and if they burn their credits on sol in 20 minutes, well, tough luck buddy, looks like you're coding by hand for the rest of the month.
Team will quickly shift. People hate losing access to ai.
Sure a Lexus is better than a used Prius, until you include price
You cant just go “oh hey, I guess they're both cars so I’m taking your lexus away, catch a cab its cheaper” and expect people to just hug you be be like “yay, thanks! I still have a job I guess! :party:”
:P
Obvious to the meanest intellect they can tell the difference between the cars.
Don't complain to me if someone responds using a stupid metaphor that proves the opposite of the point they were trying to make.
Sol already lacks judgement. It will absolutely add idiotic tests and comments. Luna is that but worse so if you account for things like going down wrong paths, producing bad results, overthinking then it could easily cost you more to get less.