DeepSeek API Pricing Update
api-docs.deepseek.com
api-docs.deepseek.com
That sounds like a policy written by someone who doesn't understand how LLM's work...
Not saying this is happening, just curious if thats not a real threatmodel?
I imagine it would be very non trivial to do it in a way that that was reliable and obfuscated enough to prevent detection for any amount of time?
It’d be much easier to hide sketchy code in an agent harness, but “vendor adds spyware to their software” isn’t a novel issue.
I think the only sort of new issue is people “allow all”ing their agents tool calls, but that’s more or less the same issue as curl | bash
Pi/OpenCode seem pretty straight foward and widely used enough for this to be viable
OMP as I understand does it's own vendoring of tools, so I assume it'd be a pain in the ass to audit, but that means you're even safe from base OS shenanigans
Hopefully this will change soon. But AI and China/US skepticism is very high. Even if the person you talk to isn't skeptic, his boss may be. And even if his boss isn't, his CFO or Legal department may use it as a political lever and therefore if you can say 'everything in europe' you dodge the tension entirely.
Yeah it's dumb.
Why use a Chinese product when a domestic or EU one is better and safer?
For European and US customers this is effectively 2x increase. I think i wll keep using both Flash and Pro as before.
EDIT: Misread numbers to believe off-peak kept old prices
Personally, I don't think we've seen the total end of dirt cheap LLMs, it's just a frontier lab doesn't want to be in business of serving half the world.
You certainly don't need Fable to code up a basic web app, any more than you need a Ferrari to go grocery shopping.
As someone from just such a country, DeepSeek 0731 was the first time I seriously started using an LLM for coding. All previous attempts were useless or ridiculously expensive.
Can't say the old prices felt "free", but it was affordable if you're careful with your cache hit rate.
The new pricing probably pushed it into the unaffordable territory for tasks where you can do without it. Probably will try opencode go if they don't also follow suite, or will have to go back to wetware.
They are hoarding HW at massive scale, they make it harder and more expensive to own
Just because you are fine with the new price doesn't mean it's not a problem
Perhaps it's time to pop this bubble
Deepseek's official API has a pretty bad privacy policy so I would assume businesses avoid them in any event
Which is a real bummer because it’s otherwise solid with excellent caching.
I'd still prefer it wasn't like that, but I guess it's a compromise ultimately, at least I'll be able to run it myself.
I know it's more complicated than that, with economics and privacy factors involved. But if I operate on the assumption that real privacy is real hard I'd rather just be guarded and careful with my prompts and whatever output I give to the LLM and expect that it will be training on that, rather than spill all my darkest secrets to some other LLM provider that pinkie-promises privacy only to leak it publicly later, anyway.
Not saying they are or will, but their privacy policy is so permissive.
Interestingly it seems that Chinese customers are even more privacy-concerned than US ones, which is why the majority of Ziphu's (GLM) business is support services to Chinese companies running their open-weight models on-prem!
If you buy Claude api access through a third party Chinese company, it's cheaper than when you buy directly.
Beside it would go against the Chinese philosophy of just using the right tool for the job. If Claude is better at a task than deepseek, they sure use Claude.
No really I don't get your claim, do you have any proof that you can source ??
Anthropic doesn't allow Claude in China and neither does OpenAI. But the rest of Asia certainly has a choice. Unless you mean socioeconomically. Also there's plenty of loopholes people in China can use to still access those models
Let's say DeepSeek is being forced to use the CANN stack, and the new pricing reflects the cost when 100% of inference is done with Huawei chips. Then, I suppose we can infer that:
* CANN stack is 1.5x~2.3x less efficient in compute
* CANN stack has 6x lower inter-connect capacity
> computer chips once again become a commodity
Ascend 950 is going for $7k to $9k with mediocre looking specs. $16k for RTX Pro 6000, $6k for RTX Pro 5000. This is not looking good.
The CEO of DeepSeek recently revealed to investors a lot about the resources available to DeepSeek, and the gap between Huawei and NVIDIA. Select quotes from the transcript (translation is a bit patchy on the source website though):
"We currently have roughly 20,000 H-equivalent compute cards"
"Huawei 950—right now Huawei gives us 16,000 cards, this should be publicly stateable."
"Like Huawei gives us roughly 16,000 cards of capacity, internet giants maybe get a hundred-something thousand, we get ten-something thousand—I think this ratio is also relatively... but this is probably just how much capacity Huawei has."
"16,000 Huawei 950 cards only equal 4,000 B-series cards."
"Huawei’s supernode, Huawei’s 950 supernode, in performance and price can completely substitute for NVIDIA’s GB200, GB300. The price is definitely more expensive, but limitedly so. Fifty percent more expensive, a hundred percent more expensive—a hundred percent more doesn’t matter, two hundred percent more doesn’t matter. For example, a hundred percent more expensive—I think it can already be considered a price-level substitute."
"I think domestic hardware might need a few years."
"I don’t quite believe that five years from now, we’ll still be stuck on the production capacity problem. Right now we’re definitely stuck on the production capacity problem—this year, next year, the year after, I think we might still be stuck on the production capacity problem, but five years later, I think maybe not necessarily—I’m still relatively optimistic."
[1] https://www.fredgao.com/p/deepseeks-liang-wenfeng-breaks-his
cache-hit 2.5x
cache-miss 1.57x
out 2.36x
# Flash, peak cache-hit 5x
cache-miss 3.14x
out 4.71x
Edit: fixed the numbers and formattingAlready outdated though I think, as GLM 5.3 is latest now :)
I am already using GLM 5.3 with their coding plan, but oddly enough the API prices don't really seem to be out yet: https://docs.z.ai/guides/overview/pricing
You'd kinda expect them to be the same as 5.2 though, seeing as that was the case with 5.1 as well (not with regular 5), but who knows.
For what it's worth, DeepSeek is still positioned as quite affordable, just not as dirt cheap as before.
(Input / Output / Cache Read, [$/M])
DeepSeek-V4-Flash:
Prev: 0.14 / 0.28 / 0.0028
Off-Peak: 0.22 (1.6x) / 0.66 (2.4x) / 0.007 (2.5x)
Peak: 0.44 (3.1x) / 1.32 (4.7x) / 0.014 (5.0x)
DeepSeek-V4-Pro:
Prev: 0.435 / 0.87 / 0.003625
Off-Peak: 0.66 (1.5x) / 1.98 (2.3x) / 0.022 (6.1x)
Peak: 1.32 (3.0x) / 3.96 (4.6x) / 0.044 (12.1x)
gpt-5.6-luna: $0.20 / $1.20 / $0.02 / $0.25 (In / Out / Cache Read / Cache Write)EDIT: formatting
EDIT2: giving up on the formatting :-/
Keep at it, I believe in you.
p.s. Thanks DSv4-Flash, for your hard work of converting a messy table into plain text.
Provider, Model Billing Input Output Cache read Cache write
DeepSeek
V4-Flash Old $0.1400 $0.2800 $0.0028 -
V4-Flash New Off-Peak $0.2200 (1.6x) $0.6600 (2.4x) $0.0070 (2.5x) -
V4-Flash New Peak $0.4400 (3.1x) $1.3200 (4.7x) $0.0140 (5.0x) -
V4-Pro Old $0.4350 $0.8700 $0.0036 -
V4-Pro New Off-Peak $0.6600 (1.5x) $1.9800 (2.3x) $0.0220 (6.1x) -
V4-Pro New Peak $1.3200 (3.0x) $3.9600 (4.6x) $0.0440 (12.1x) -
OpenAI
GPT-5.6 Sol $5.0000 $30.000 $0.5000 $6.2500
GPT-5.6 Terra $2.0000 $12.000 $0.2000 $2.5000
GPT-5.6 Luna $0.2000 $1.2000 $0.0200 $0.2500
Anthropic
Claude Fable 5 $10.000 $50.000 $1.0000 $12.500
Claude Opus 5 $5.0000 $25.000 $0.5000 $6.2500
Claude Sonnet 5 $2.0000 $10.000 $0.2000 $2.5000
Moonshot
Kimi K3 $3.0000 $15.000 $0.3000 -
Z.AI
GLM 5.2 $1.4000 $4.4000 $0.2600 -I’m curious about how openrouter and Luna prices will change in response.
DeepSeek was hugely underpricing cache hit pricing before and even after this increase they're still cheaper on that metric than every other provider I'm aware of, but it will put an end to those "I used 1 billion tokens and spent $4" reports.
I think that Flash is still a usable model but Pro is DOA... Even before the price difference between Flash and Pro, vs the intelligence / problem solving / tool calling did not make sense. But now that gap has widen even more. And there are just too many competitors models now close to that Pro price range.
Especially when we compare that competitive models offer subscription services that easily cut down the token price by 1:10. That makes Pro especially a bad value.
We shall see what the 3th party market is going to do, but i suspect that prices will be increased. If the argument was that DeepSeek increases price as they lack capacity, a company with access to billions, other 3th party providers that need to rent and have less optimized infrastructures will increase prices. Especially if they get hit hard with people moving around.
Its like we always see the same issue with popular models.
* GLM 5.2 is good, capacity issues, API price up, subscription heavy nerfs. * Kimi K3 is good, capacity issues, API price up, subscription heavy nerfs. * DeepSeek V4 GA is good, capacity issues, API price up * OpenAI GLM 5m, 10m active users. Subscription usage is sneakily tightened more and more. * Anthropic Opus too popular, ...
That is the main issue. The AI users are people who actively easily move between companies. Pushing peak loads to each unprepared company, releasing load on the "less desired". And round we go ...
The benchmarks show that Luna is significantly faster, but I think those are very complex tasks for which you'd probably want a bigger model anyway. (e.g. Sol is much faster than Luna at the same tasks.)
So I'm wondering if there's any difference for smaller tasks, or if they're basically matched now.
https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
If I'm reading the benchmarks right, they now went from being much cheaper than Luna (but twice as slow), to being roughly same price (but twice as slow).
So all else being equal, where I would previously have used DeepSeek, I can just use Luna, and get the same result twice as fast?
(Yeah I know benchmarks are mostly nonsense, but the ones measuring time are real, and it's the most precious resource.)
Also if you're considering Luna, I assume you don't care about this but I think it's worth pointing out: a major advantage of DS is the ability to self-host or choose a different host. As a customer that gives you much more negotiating power and potential privacy guarantees.
It’s a race to the bottom, and the bottom is unlimited use for a flat monthly rate.
Granular pricing (tokens, minutes, etc) is pretty anti-customer generates less revenue than customer value-based subscriptions (why SaaS is such a good business model)
Consumers don’t generally get usage-based pricing because of the inconvenience and unpredictability, but B2B SaaS products utilize usage-based pricing all the time.
Pricing software is a game of estimating both software value and the purchasing power for customers. Only the latter might have any available data and even then it won’t be sliced the right way for any in depth statistical analysis that an actuary would perform to underwrite risk.
It’s much more traditionally a more salesperson like background where being in the target market or having strong connections to it dominates efficacy.
More on my approach here: Https://forstarters.substack.com
But my Claude Max subscription? If I have any of my limit left the day of my reset, I’ll go and fire off research workflows with a bunch of parallel agents to explore whatever dumb ideas I had the past week. And there’s a 50:50 chance I’ll forget about it and never read the output.
I am however going to fire off a half assed prompt when the marginal cost is zero, even if I don’t use it (which is par for the course, I probably throw out two thirds of anything the AI writes anyway be it code or prose).
Any experience with it?
that's a good outcome - it means they're fungible, and easily available.
The summary of this paper describes my sentiment in better words than I have:
https://www.nature.com/articles/s41599-025-05868-8
It’s very easy for the average person to mistake linguistic ability and simulated problem solving for intelligence and sentience.
I do not believe current AI or LLMs are conscious, but there is no proof one way or another that they can or cannot be. The paper authors are making up their own definitions and building an argument from them
I, and apparently many others, don’t think it would be any useful to describe the mathematical properties of an AI as consciousness. To me it is inherently a way to describe the “experience” arising from physical processes in biological beings as ourselves.
That’s what the argument comes down to for me. Could an LLM “fall unconscious”?
Consciousness is defined in the human context. We can just say “hey buddy that’s not biological enough to qualify.”
Though of course you're talking about data centers, and romanticizing them rather than the AI itself.
There are various theories around model collapse when you train on too much AI generated data (that's not for distillation).
Right now, they give 4100 credits for Luna and 63 000 for Deepseek on their prepaid plan (both are 2x)
Do they just set a super low caching time and hope that drops effective cache rates low enough? Do all other providers somehow overcharge by that much? Are they just going to sell it as a loss leader?
This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead.
From a consumer viewpoint a more interesting metric than the raw costs is
cached cost * hitrate + input cost * (1 - hitrate)
from a personal standpoint rather than a per-provider one (e.g. if OpenRouter is blindly dispatching your requests you might have a bad time).I don't have a clue on what the real cost to inference providers comes out to, but it seems really weird that there would be such a big gap, in what should be a pretty competitive market.
This can somewhat be the case, depending on your config. I updated mine to make DeepSeek high priority because I was having a lot of cache misses and reliability issues with the default (cheapest (at face value)) providers, and cost was actually higher overall than anticipated. Was smooth sailing from then; might have to tweak things again now pricing has changed though.
Not sure if this was an OpenRouter issue or with the other inference providers.
I pasted the same prompt into OpenCode, set to Deepseek v4 flash free and did it first try.
I'm was going to purchase Opencode GO to try it, but seems my timing is really bad :( hope it doesn't go up too much in Opencode or they find other providers. Bad timing!
But DeepSeek v4 Pro is a far more capable model and still cheaper than anything that it competes with, from what I can see.
Thank you in advance!
I am a person that buys into a tool or a process and expects it to be part of the life with no major changes through the years (or as long as the need exists). But AI? You buy into something today, not 2 weeks have passed and there's already a large "update" introduced to the conditions or the optimal usage patterns you should be adopting.
It's tiring. Makes all prices and offers feel so unreliable and gets me a bit more disinterested each time they change.
These being open, you can keep using the old models indefinitely for as long as there are providers offering them.
Fully knowing that it is a new industry living its own infancy, it is perfectly normal that there is instability and numerous swings on pricing, conditions, or direction.
But it's not less real that such process can produce churn and consumer fatigue.
The same motivation of course - the GPUs have a finite service lifetime, so to maximize revenue you need to keep them busy 24x7.
https://www.bloomberg.com/news/articles/2026-06-17/microsoft...
Bytedance which runs China’s most popular Doubao AI chatbot; is spending $70B in CapEx this year, most of it outside of China (Malaysia, Thailand, Brazil, etc; and they are allowed to lease NVIDIA chips). This is roughly 50% of Microsoft CapEx.
https://www.tomshardware.com/pc-components/gpus/chinas-byted...
The big Covid migrations (startup prople migrating to the countryside),
Will we see the big AI migrations (people travelling to where AI is the cheapest)?
https://www.baseten.co/pricing/
If anyone has tried Baseten versions of these Chinese frontier models, let me know what you found.
If you check on OpenRouter, some other providers serve V4 Flash at seemingly cheaper normal input/output tokens rates, but with a huge caveat: they have at least a 5x increase of the cache hit cost of the official API, some have a 10x+. No provider comes close to Deepseek's old low cache prices, and cache is 90%+ of what matters in agentic sessions.
Closest comparison:
- Deepseek: $0.14/$0.28 with $0.0028 cache hit cost for official API
- DeepInfra: $0.08/$0.18 (cheaper base rates!) with $0.016 cache hit (almost 6x!! Deepseek's current cache cost)
Another great example is Kimi K3, official API is $3/$15 and the cheapest provider on OpenRouter is $2.8/$14, only a tiny difference.
DeepSeek-V4-Flash (off-peak, x2 for peak)
* Cache Hit $0.007 (x2.5)
* Cache Miss $0.22 (x1.5)
* Output $0.66 (x2.25)
DeepSeek-V4-Pro (off-peak, x2 for peak)
* Cache Hit $0.022 (x6)
* Cache Miss $0.66 (x1.5)
* Output $1.98 (x2.25)
Peak Hours: 01:00–04:00 and 06:00–10:00 UTC
Effective from: 16:00, August 16, 2026 (UTC)
- but what matter - is cache hit
even now deepseek's off-peak hours for cache hit (0.007) is lower than other providers (~0.01)
It’s still cheaper than everybody else.