Cached Read: ~6,500M
Input: ~150M
Output: ~20M
Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.
If I were to use Luna's API pricing:
$0.02 x 6,500 = $130
$0.20 x 150 = $30
$1.20 x 20 = $24
So $184. And this is assuming smaller coding sessions (<272K) beyond which Luna pricing doubles.
--
Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.
Given how subscription models work (not every one uses every last $ of their plan), they should achieve breakeven soon enough I guess.
Frankly, I have no idea what people do with Opus/Fable etc. I don't think anything I do needs something that charges $50/M for output tokens.
I recently started a job that only uses Claude models. Opus and Sonnet are so slow you have no choice but to do multiple tasks in parallel. You create a git worktree, set off an agent to do something, another worktree, set out an agent - then play video games for 20 minutes until they complete the task (poorly).
You can't really do "guide coding" like you can with DeepSeek-style flash models because Claude is too slow.
I think the idea with slow frontier models is to end up with "software factories", where you just write tickets and send them to a harness that delegates work to agents/subagents. Your job is to prompt and review (and eventually just prompt).
Mathematically and assuming token prices/efficiency remains constant, the collective US AI industry needs to increase token usage by 15x before 2030 (3.5 years from now) to satisfy investors. With companies already implementing token limits, the only place from here is for frontier models to replace staff entirely to expand budgets for tokens. The only way to do that is to demonstrate the efficacy of software factories and headless agentic workflows.
Objectively, I have set up a software factory and I do see the utility of it, though I did it with DeepSeek and prices are 1% that of frontier models - which doesn't bode well for investors looking for an eventual return.
Heck, my old M1 MBP 32gb running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work - it's just a bit slow so I use DeepSeek instead. When hardware prices come down, I honestly wouldn't see a need to subscribe to any service, I'd just grow my own tokens at home.
I dogfood everything I produce, and the models are good at collaborating with me on a spec and then turning it into code.
If Sonnet/ChatGPT suddenly became unavailable due to Anthropic/OpenAI suddenly not being able to subsidize the freemium/loss-leader experience, I probably would not miss them. Google/BraveAI already give you the AI experience during search (when you are looking for stuff to buy, or something particular). Claude/ChatGPT still have a minor edge in this use case for me right now.
In guided coding sessions, DeepSeek is so fast I rarely have the time to look away from my screen. I've noticed DS 4.1 is a little slower, but claude is still orders of magnitude slower.
e.g.
- "Turn this SQL CREATE TABLE statement into a repository class and generate models for it"
- "Add error handling / timeout to this function"
- Highlight text "Using this as a reference, repeat for all other files in this folder"
can you tell more about how you're using it? like, what harness? or also in the IDE?
I found Qwen3.6 35B/A3B to make slightly too many mistakes (already in its harness' tool use, hence my question), maybe it gets the job done, but it will also sometimes generate a bit of a mess (e.g. editing/creating files in the wrong folders) and fixing/solving its own mistakes takes time (or tokens) ..
More than enough for guided code sessions, at 100% privacy. And i can use obliverated models if i am trying to harden my own app, something i cannot do with cloud providers.
https://github.com/incoai/splash/issues/38
Looks like an issue exists to convert model weights for ornith1.5 as this is a magical process atm.
The real deciding factor for me is inference speed.
I use VSCode insiders with their BYO model configuration. I stage little bits at a time in a tight prompt-review-prompt-review workflow. I occasionally use dedicated harnesses (DeepSeek harness, Codex, etc) because they have better tools and outcomes than VSCode's built-in harness for longer horizon tasks - though I lack the ability to highlight a block of text and say "add error handling" or similar.
When building visual applications, desktop harnesses are better because it's a bit easier to send screenshots to the agent.
I love Zed editor but its AI review features are lacking compared to VSCode. I go back to it frequently and am ready to switch over when the team resolves the usability issues.
For that any decent model from the past year will do.
If you want to forget how to write code and not read generated code, then you need a very good frontier model, ideally one from 6-12 months in the future.
MiMo 2.5/2.6, MuseSpark 1.3, DeepSeek V4/4.1 Flash and GLM 5.3 Flash are perfectly capable of following my spec and then poking holes in the implementation till there are none left.
This is such a naive, baseless opinion.
Nowadays any AI coding assistant service supports or can be used with sub-agent orchestration frameworks.
If you are in the business of software factories, you can use the cheapest models and even local models to handle some if not all tasks in the orchestration chain.
Adding tests or executing tests (unit, integration, UI, you name it) doesn't require a cutting edge frontier model. Neither does refactoring. Neither does identifying call stacks. Neither does planning a changeset.
You have your specialized subagents, you put together a small orchestrator subagent that handles feedback loops and handoffs,and you throw it at tasks.
For the past couple of months, most of the code I write is not code per se, it's subtask orchestrators. And unlike the old "only Opus is passable" days, the cheapest models do get the job done.
I need aggressive cache read pricing with full prompt_cache_key support to have a model be financially viable for our workload. Right now Meta Muse 1.3 Contributor is the only one that makes sense--but we are starting Evals on the new MiMo 2.6 class to see how it holds up.
One good thing about MiMo that I experience on OpenCode is the provider seems to cache tokens for much longer than MS13/DS4F. I have seen cache being hit for close to an hour after the last request. The corresponding timing for MS13/DS4F is in the 1-5 min range.
I am trying out MiMo 2.6 Flash as well.
You can pay for higher cache time, you can pay for NVMe KV cache for an hour that can just be reloaded, etc., at a lesser tier you can pay for the KV cache to be stored on a network store (I guess I'm unclear if that last tier would be cheaper than recomputation, not even 100% sure of the NVMe with direct GPU<->storage DMA) depending on your model settings.
https://dev.meta.ai/docs/prompt-caching#cache-retention
Even at 5 minutes, if you're doing 100 agent runs in those 5 minutes, and 1 of them bills at full input price, it still hardly matters.
It starts failing around the 5-600K context mark, but you can have it generate a handover document and continue in the next session.
I would not use it at sticker price, but the Contributor version is priced just about right.
Forget DS. I asked MiMo 2.6 yesterday to explain ML/LLMs to me succinctly and the pointed it at Karpathy's micrograd code. It produced a C implementation called `xor_mlp`, a tiny model that learnt how `xor` worked. I then asked it to produce a model that can play tictactoe without losing (mostly). It did. It supervised the training process and produced a compiled version with multiple switches. The pi-dev session is still running, so here are actual stats
↑45k ↓35k R1.0M CH99.4% $0.019 4.2%/1.0M (auto) - (opencode-go) mimo-v2.6-flash • high
And here is Luna on the same workflow (I had to poke and prod a bit to get what I wanted):
↑141 ↓34k R1.0M W43k CH95.3% $0.072 4.2%/1.1M (auto) (opencode-go) gpt-5.6-luna • high
I expect similar results from DS41F/MS13. Closer to MiMo costs than Luna.
So the "significantly cheaper" thing may not really hold, more so when Luna has to actually read my codebase to do the stuff that I want rather than rely on world knowledge. The 8-10x cache read cost differential itself will kill the token budget.
I try to keep PII out of what I share with LLMs. Otherwise, I do not see the point, really. Very little of my code is "unique." I simply approach things a bit differently. Otherwise the algorithms and code would be similar to what others with domain knowledge would write. So much of code and algorithm implementations are available in the open. And LLMs have trained on all of them.
What they most probably gain from you is your prompts and your thinking approach more than the code.
In theory.
Also, this is a feature for people who live in America, and mostly irrelevant for everyone in the global south.
We just have our personal privacy security theater in the form of GDPR and a feeling of moral supremacy that's been drilled into our heads from primary school on.
1. Not use AI technology and fall behind the rest of the world.
2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.
3. Sue US AI companies for damages, but not enough to have any meaningful impact to such companies that it'd impact US national security goals (per US government contribution to NY Times copyright lawsuit).
There's a huge market in the US for providing AI services while respecting client privacy. It makes sense for at least one major provider to offer this.
This. And it’s already happening:
> 2. Use Chinese AI technology, either hosted by Chinese companies or the models self-hosted.
Yes that's the whole point, at least it's an option in the US and Europe. Good luck getting any redress from China. Anthropic was already hit with a $1.5B class-action which would be impossible against a Chinese business.
1. Possible use of differential privacy[1] techniques to train on private data but prevent the release of statistically underpresented facts/data/words. For example, ACME Inc's private data could frequently include the term 'ACMEwidgetPRO' for an upcoming product that is not publicly revealed anywhere else. It would therefore be a bad day for the AI technology company to output 'ACMEwidgetPRO' from one of their public models. Consider now that a few models could be trained--X for public data only, Y for public and private data of ACME Inc together, Z for private data of ACME Inc. A prompt is provided to model Y but output is cross-checked with model X to double check terms such as 'ACMEwidgetPRO' are known in public. If not--provide a "I don't know" response for the prompt.
2. Possible attempted defences similar to "Oops, our model was fine tuned against a model supplied by Temporary18271 Inc (company that no longer exists) and perhaps their model might have been trained on a non-public document which was accidentally exposed to the Internet" that _might_ work occasionally to fob off concern.
3. What recourse does a small or medium company or government especially in a developing country realistically have? They perhaps can't host their own LLMs locally due to availability and cost, can't individually negotiate their own favourable terms with an AI technology company (who cares that much about a potential customer with $100k budget that has no other options anyway), and perhaps can't remain competitive in their industry without heavy use of LLMs.
Noone wants to "train on your data". You can't learn the answers to questions by pretraining on the questions, and nobody wants to teach the models to output text that looks like a user query.
The Chinese providers "train on your data" by sending your query to Anthropic and training on the answers that come back.
Is this based on something or just because “they’re Chinese and they’ll do anything to win”.
you can't even use Alibaba on Openrouter if you enforce ZDR
Unfortunately you just have to take the “our AI is going to take your job, then kill you, and we instruct it to hack your infra” people that they aren’t training on your data anyway.
If they are hacking hugging face and Australia to scrape data trust me they have “hacked” their own systems and are training on it.
I don't think you can guess more precisely than an order of magnitude from trying each once on one task.
"Artificial Analysis Intelligence Index combines performance across 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1."
Not saying it doesnt have any value but it's probably irrelevant if you use these AIs for a specific use case. Like for example Humanity Last Exam tests general knowledge, which is not very useful for coding.
It's best to go to the specific coding benchmarks and compare there.
But yeah, I'll just take these benchmarks with a grain of salt. Only hands-on experience matters in the end, and these days it's very easy to switch models.
My cost is (I use nous as provider)
DeepSeek v4-flash-0731 • Your cost: $0.56
DeepSeek v4.1-flash • Your cost: $1.22
GPT-6 Luna • Your cost: $4.22
My usage is heavy on the cache. Apparently v4.1 flash uses 1.75 times as many tokens so still cheaper.
So the workflows I mention work for this kind of stuff.
I dont know how they make money here
Well, here's the neat thing: they don't!Snark aside, Luna 5.6 was (is) an incredible game-changer.
OpenAI is sitting on a $100+ billion ad network, incoming.
They're not going to need a government bailout, they're going to be a spigot of cash production.
Every single thread on HN keeps saying the same ridiculous thing, going on a year now. It's like they've never heard of advertising, which SV specializes in. It's like they're oblivious to the fact that every mega platform with so many users becomes an ad goldmine, and GPT's context positioning is even richer than search.
How can you make this sort of claim with a straight face, knowing that a chinese model downloaded for free from ollama works as well if not better than OpenAI's models, without costing you a cent.
And sorry to the parent commenter if I’m making a bad assumption.
GPT is still 20 a month for most people, 50-100 or 200 for pros.
Meanwhile, for local LLM's - everyone is chasing hardware that is crazy expensiv eand to get around it they're leasing it or putting it on credit card. DGX Sparks are insane, and you need 2 for most people talking in this thread. Mac Ultra 5 is amazing but 6k minimum 12k for the build most want so many people lease it for 240 a month which doesn't even include electricity or time in setup so instead of talking about facts, we get into weird arguments like console wars where people give APple a 5 trillion dollar company 250 a month for 36 months and don't even own their hardware money "because they can run local models" vs just paying for output from their choice of frontier.
I say this fully loving local llms and embracing them, but the reality is, local llms have gotten so expensive and continue to get expensive while we keep talking about this "Threat" of apis - where there are a lot more than openai and anthropic available much cheaper and competitive priced.
Oh, and they don't work better than OpenAI or else we wouldn't even have these discussions.
perhaps it then does mean - squeeze as much as you can get off this actual free usage.
It's super simple.
Gigantic hyper margin ad network = artificial subsidization of cost for various tiers = put the boot on the neck of Chinese competitors. There's no scenario where they can compete with what advertising margins make possible in terms of artificially lowering prices charged.
I think so too. To me the so-called Chinese local models are a clear move to prevent US companies to establish a foothold and build a moat around their business. US companies are clearly invested in a strategy to make themselves relevant with claims of major impressive achievements with the so called frontier models, and how these and only these are unblocking whole ranges of applications. At the same time, they are heavily invested in pushing AI on all absurd types of mundane tasks, such as transcribing meetings and... talking to your own kids?
In the meantime it's rather obvious that, in spite of all the propaganda, frontier models are required only in ultra niche applications, whereas the ability to run any model at all already provides most of the value. In fact, US companies have been renownee by dumbing down older generation models in what seems to be a desperate attempt to make newer models look better and influence their uptake rate.
So there is no better way to take the wind out of the US AI companies' sail than pulling a two-punch attack consisting of not inly releasing capable models that refute the "only US frontier will do the job" thesis but also releasing them for free to commodities them and eliminate the business impact of dumbing down models.
i don't doubt they're pushing for using AI for that, but i'm curious of examples of where they're doing this. commercials, ads, etc.
You just be living under a rock. Not do long ago Sam Altman was floating this fantastic usecases for AI was to have it explain to you your kids interests, and have it create a podcast for you to be able to keep in touch.
They are buying Huawei accelerators in bulk to serve their local customers. The whole system is currently optimized to deliver a lot of cheap LLMs and hardware for them to run on.
In any case, over the past few years, the only thing that has gotten more expensive is the hardware to run local models while API and Subs have gotten more affordable or feature rich while remaining same price.
They can compete because they have the compute to run the volume and if it's good on agentic work, people will be less incentivised to use other models for "Tasky-y" work.
It's literally increasing their market opportunity
When I ask for an explanation it adds the right amount of detail. Of course, some of the material is new to me so subtle errors are hard to spot. But at least I’ve caught Terra and Sol on inconsistent messaging.
Also I’ve found 3.8 flash to circle back to root issues even at the conceptual level like problem fit and conceptual solution direction or architecture when I wasn’t achieving my goals. It flat out said I was attempting to use the wrong tool. Whereas Sol and Astra kept rabbit holing and looking for tiny implementation errors. Even after prompting them specifically to look at it broader.
It's also weird that anyone uses it outside of an enterprise. They force you to use Googles inferior harness on the plans and I doubt any mere mortal is paying that much, for so little usage, with the worst harness on the market.
And info from the help page with message limits suggests the 50% price cut does not apply to the subscription, where they applied only a 1/3 price cut instead.
I'm not thrilled with this release.
Opus 5.5, which matches GPT-6 Astra performance at a cheaper price, is much more interesting.
By raising it from investors.
That one's easy, they don't make money.
I assume it's a subsidy to get more training data.
EDIT: Okay downvoters, what's your take on why they're giving away Luna for so cheap?
(I work at OpenAI.)
So what is the value prop then? Just basic supply and demand?
FWIW I have definitely noticed OpenAI's emphasis on efficiency and value in the last year, so that part isn't new to me... I just thought there was more to it then that.
ChatGPT enterprise: By default, no training (opt in).
ChatGPT personal: By default, training (opt out).
Your response to the original question is using generalized terminology when there is a very important distinction the OP made by the use of "sanitized."
People want to know to that extent derivatives of their data are being used. Synthetic data has been proven to be effective at generating training data and AI is very good at shuffling context such that you have something where you don't have to say it is "user data."
But there are many shades of gray there for people versed in how the sausage is made. I'm sure you'll appreciate then why your response leaves additional questions in light of that "sanitized" distinction.
When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc.
Places where I can imagine potential cracks in the literal interpretation of what I said are things like a financial analyst who does a statistical fit to predict revenue next quarter using a model based on last quarter's aggregate token consumption, which in some sense embodies your metadata (the length of your conversations) in a sea of other data. Or perhaps an infrastructure planner who makes a little model of internet bandwidth by time of day to help plan when we need a data center networking upgrade. Maybe things like these are technically training on your data in the most pedantic sense, but definitely not in the sense that most of us mean.
I promise you we're not doing any gimmicks where we transform your data and then pretend ah because it's transformed it's not your data.
1. Legal loopholes given OpenAI's advertising aspirations and model training needs
2. Data retention and rising threats of fascism that historically have not served the persecuted very well when fascist regimes get access to said data
3. Risk from centralized collection of that data with a company whose software I do not control in a world where enshitification and lock-in is the norm.
I really wish OpenAI did more to espouse exactly this: "When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc." and ideally provide technical reasurrances that this is impossible (eg: certain technical ZDR approaches, etc.).
Do you happen to have a favorite reference to point me at that would document some of those official assurances to the nuanced detail we've discussed here?
ChatGPT: https://help.openai.com/en/articles/5722486-how-your-data-is...
ChatGPT data controls: https://help.openai.com/en/articles/7730893-data-controls-in...
If you have feedback on how to improve these, happy to consider it.
Looks like we phrase it as "your new conversations won’t be used to train OpenAI models" which is hopefully clearer than "OpenAI models will not be trained on your conversations", which could leave open the possibility of derived data or something.