https://artificialanalysis.ai/models/comparisons/gpt-6-luna-...
GPT 6 Luna closed the gap significantly for sure (it seems to be about twice as expensive as DS v4.1 Flash), but Deepseek v4.1 Flash is still the best value model and capable enough for almost everything I need to do. Sometimes if it's babbling or can't nail down a solution I switch to Sol for one prompt, get the solution, then switch back to DS. I use 4.1 Flash almost exclusively though for both planning and implementation these days.
I was a heavy Kimi k2.5/2.6 user but since 2.7 Kimi has gone way downhill -- even the previous models. I think they got under heavy load and had to quantise their models to avoid going broke.
Is a zero-shot, zero-context prompt a useful benchmark? Yes, in the absolute sense. Does it reflect how teams would use it in the real world? I think in real-world use cases (IME), DeepSeek gets the job done.
The DS one used Claude Code via OpenRouter, the Luna one used Codex. I'd say a big difference in cost is coming from the harness, and quite possibly the different meanings of "one shot" in each of those harnesses. The Luna one probably spent less money, but might have also done far less real browser testing and testing/verification is the expensive part.
The better graphics is probably just Codex system prompt.
In terms of real world coding usage, I am finding that DS v4.1 Flash is about half the price of Luna for comparable workloads. G6L is super cheap for sure, and super capable. It's by far the best coding model from a frontier lab for everyday coding work, but DS v4.1 is even cheaper and no less capable in my experience.