On the hardware side you can run some benchmarks on the hardware (or use other people's benchmarks) and get an idea of the tokens/second you can get from the machine. Normalize this for your usage pattern (and do your best to implement batch processing where you are able to, which will save you money on both methods) and you have a basic idea of how much it would cost per token.
Then you compare that to the cost of something like GPT5, which is a bit simpler because the cost per (million) token is something you can grab off of a website.
You'd be surprised how much money running something like DeepSeek (or if you prefer a more established company, Qwen3) will save you over the cloud systems.
That's just one factor though. Another is what hardware you can actually run things on. DeepSeek and Qwen will function on cheap GPUs that other models will simply choke on.
Or with somebody else's.
If you don't have strict data residency requirements, and if you aren't doing this at an extremely large scale, doing it on somebody else's hardware makes much more economic sense.
If you use MoE models (al modern >70B models are MoE), GPU utilization increases with batch size. If you don't have enough requests to keep GPUs properly fed 24/7, those GPUs will end up underutilized.
Sometimes underutilization is okay, if your system needs to be airgapped for example, but that's not an economics discussion any more.
Unlike e.g. video streaming workloads, LLMs can be hosted on the other side of the world from where the user is, and the difference is barely going to be noticeable. This means you can keep GPUs fed by bringing in workloads from other timezones when your cluster would otherwise be idle. Unless you're a large, worldwide organization, that is difficult to do if you're using your own hardware.
Isn't that true for any LLM, MoE or not? In fact, doesn't that apply to most concepts within ML, as long as it's possible to do batching at all, you can scale it up and utilize more of the GPU, until you saturate some part of the process.
What's cheap nowdays? I'm out of the loop. Does anything ever run on integrated AMD that is Ryzen AI that comes in framework motherboards? Is under 1k americans cheap?
[1] https://youtube.com/@digitalspaceport?si=NrZL7MNu80vvAshx
When used with crush/opencode they are close to Claude performance.
Nothing that runs on a 4090 would compete but Deepseek on openrouter is still 25x cheaper than claude
Is it? Or only when you don’t factor in Claude cached context? I’ve consistently found it pointless to use open models because the price of the good ones is so close to cached context on Claude that I don’t need them.
Things get a lot more easier at lower quantisation, higher parameter space, and there's a lot of people's whose jobs for AI are "Extract sentiment from text" or "bin into one of these 5 categories" where that's probably fine.
And without specifying your quantization level it's hard to know what you mean by "not usable"
Anyway if you really wanted to try cheap distilled/quantized models locally you would be using used v100 Teslas and not 4 year old single chip gaming GPUs.
Uh, Deepseek will not (unless you are referring to one of their older R1 finetuned variants). But any flagship Deepseek model will require 16x A100/H100+ with NVL in FP8.
We can't judge on training cost, that's true.
You are right that we can directly observe the cost of inference for open models.
A few days ago I read an article saying the Chinese utilities have a pricing structure that favors high-tech industries (say, an AI data center), making the difference by charging more the energy-intensive but less sophisticated industries (an aluminium smelter, for example).
Admittedly, there are some advantages when you do central and long-term economic planning.
That's how it supposed to work.
I hope it does, though I'm doubtful because distribution is important. You can't beat "ChatGPT" as a brand in laypeople's minds (unless perhaps you give them a massive "Temu: Shop Like A Billionaire" commercial campaign).
Closed source AI is almost by design morphing into an industrial, infrastructure-heavy rocket science that commoners can't keep up with. The companies pushing it are building an industry we can't participate or share in. They're cordoning off areas of tech and staking ground for themselves. It's placing a steep fence around tech.
I hope every such closed source AI effort is met with equivalent open source and that the investments made into closed AI go to zero.
The most likely outcome is that Google, OpenAI, and Anthropic win and every other "lab"-shaped company dies an expensive death. RunwayML spent hundreds of millions and they're barely noticeable now.
These open source models hasten the deaths of the second tier also-ran companies. As much as I hope for dents in the big three, I'm doubtful.
Even when the technical people understood that, it would be too much of a political quagmire within their company when it became known to the higher ups. It just isn’t worth the political capital.
They would feel the same way about using xAI or maybe even Facebook models.
https://sg.finance.yahoo.com/news/airbnb-picks-alibabas-qwen...
The fact that it's customer service means it's dealing with text entered by customers, which has privacy and other consequences.
So no, it's not "pretty inconsequential". Many more companies fit a profile like that than whatever arbitrary criteria you might have in mind for "consequential".
2020 - I was a mid level (L5) cloud consultant at AWS with only two years of total AWS experience and that was only at a small startup before then. Yet every customer took my (what in hindsight might not have been the best) advice all of the time without questioning it as long as it met their business goals. Just because I had @amazon.com as my email address.
Late 2023 - I was the subject matter expert in a niche of a niche in AWS that the customer focused on and it was still almost impossible to get someone to listen to a consultant from a shitty third rate consulting company.
2025 - I left the shitty consulting company last year after only a year and now work for one with a much better reputation and I have a better title “staff consultant”. I also play the game and be sure to mention that I’m former “AWS ProServe” when I’m doing introductions. Now people listen to me again.
All tech companies offering free services.
I’m not saying this to insult the technical capabilities of Uber. But it doesn’t have the economics that most tech companies have - high fixed costs and very low marginal costs. Uber has high marginal costs saving a little on inference isn’t going to make a difference.
Obviously, some US brands do not compete on price, but other than maybe Jeep and Tesla, those have a small market penetration.
All the clouds compete on price. Do you really think it is that differentiated? Google, Amazon and Microsoft all offer special deals to sign big companies up and globally too.
Microsoft doesn’t compete on price. Their major competitive advantage is Big Enterprise is already big into Microsoft and it’s much easier to get them to come onto Azure. They compete on price only when it comes to making Windows workloads Bd SQL Server cheaper than running on other providers.
AWS is the default choice for legacy reasons and it definitely has services an offerings that Google doesn’t have. I have never once been on a sales call where the sales person emphasizes that AWS is cheaper.
As far as GCP, they are so bad at evterprise sales, we never really looked at them as serious competition.
Sure AWS will throw credits in for migrations and professional services both internally and for third party partners. But no CFO is going to look at just the short term credits.
Despite all that and whatever you say, the fact is you do compete. It doesn't have to be a race to the bottom.
So Cloudfront free tier and the latest discount bundles etc aren't to compete? People have also negotiated private pricing way below list price and a lot cheaper than competitors.
Similarly was the Dynamodb price cuts not due to competition?
I can give way more examples...
All technology gets cheaper over time. There is a difference between lowering price in response to competitors and finding the profit maximizing price based on supply and demand.
AWS was lowering prices to increase demand before GCP and Azure were a thing.
Jassy said right before he became CEO of Amazon and he was still over AWS that only 5% of IT spend was on any cloud provider. They are capturing non consumption and marketing value of AWS vs that.
While I don’t have any insider experience about Azure, looking on the outside, I would think that Azure’s go to market is also not competing against AWS on price, but trying to get on prem customers on Azure.
And most startups are just doing prompt engineering that will never go anywhere. The big companies will just throw a couple of developers at the feature and add it to their existing business.
Before that I spent 6 years working between 3 companies in health care in a tech lead role. I’m 100% sure that any of those companies would I have immediately questioned my judgment for suggesting DeepSeek if had been a thing.
Absolutely none of them would ever have touched DeepSeek.
If you'd spent anytime working at one for swe you won't have access to popular open source frameworks, let alone Chinese LLMs. The LLM development is mostly occurring through collaborations with the regional LLM businesses or internal labs.
https://www.ecfr.gov/current/title-17/chapter-II/part-240/su...
Note: I am neither a lawyer nor in financial circles, but I do have an interest in the effects of market design and regulation as we get into a more deeply automated space.
https://docs.aws.amazon.com/sagemaker/latest/dg/clarify-mode...
Of course you’ll always have exceptions (government, military and etc.), but for private, winner will take it all.
Any kind of hardware that is somehow connected to the wired or wireless communication interfaces is much more dangerous than any software.
Backdoors embedded in such hardware devices may be impossible to identify before being activated by the reception of some "magic" signals from outside.
Companies just need to get to the “if” part first. That or they wash their hand by using a reseller that can use whatever it wants under the hood.
Although I did just check what regions AWS bedrock support Deepseek and their govcloud regions do not, so that's a good reason not to use it. Still, on prem on a segmented network, following CMMC, probably permissable
Chinese models generally aren't but DeepSeek specifically is at this point.
Well for non-American companies, you have the choice between Chinese models that don't send data home, and American ones that do, with both countries being more or less equally threatening.
I think if Mistral can just stay close enough to the race it will win many customers by not doing anything.
I'm not sure if technical people who don't understand this deserve the moniker technical in this context.
Nobody is winning in this area until these things run in full on single graphics cards. Which is sufficient compute to run even most of the complex tasks.
You already have agents, that can do a lot of "thinking", which is just generating guided context, then using that context to do tasks.
You already have Vector Databases that are used as context stores with information retrieval.
Fundamentally, you can have the same exact performance on a lot of task whether all the information exists in the model, or you use a smaller model with a bunch of context around it for guidance.
So instead of wasting energy and time encoding the knowledge information into the model, making the size large, you could have an "agent-first" model along with just files of vector databases, and the model can fit in a single graphics cards, take the question, decide which vector db it wants to load, and then essentially answer the question in the same way. At $50 per TB from SSD not only do you gain massive cost efficiency, but you also gain the ability to run a lot more inference cheaper, which can be used for refining things, background processing, and so on.
In any case, models are useful, even when they don't hit these efficiency targets you are projecting. Just like cars are useful, even when they are bigger than a pack of cards.
Its also not a matter of it working or not. It already works. Take a small model that fits on a GPU with a large context window, like Gemma 27b or smaller ones, give it a whole bunch of context on the topic, and ask it questions and it will generate very accurate results based on the context.
So instead of encoding everything into the model itself, you can just take training data, store it in vector DBs, and train a model to retrieve that data based on query, and then the rest of it is just training context extraction.
Oh, be more creative. One simple way to make money off your idea is:
(1) Get a hedge fund to finance your R&D.
(2) Hedge fund shorts AI cloud providers and other relevant companies.
(3) Your R&D pans out and the AI cloud providers' stock tanks.
(4) The hedge fund makes a profit.
Though I don't understand: wouldn't your idea work work when served from the cloud, too? If what you are saying is true, you'd provide a better service at lower cost?
However the issue with "funding" isn't as simple as that statement above. Remember, modern funding is not about value its about hype. There is a reason why CEOs like Jenson say that if they could go back in time, they would never start their companies knowing the bullshit they have to walk through.
Ive also had my fair share of experiences in trying to get startups off the ground - for example, back around 2018, I was working on a system that would take your existing AWS cloud setup, and move it all to EC2s with self hosted services, which saved people money in the long run. I had proof of concept working and everything. The issue that I ran into when trying to get funding to build this out into a full blown product/service that I didn't realize is that being on AWS services for companies was equivalent to a person wearing an expensive business suit to a sales meeting - it was fact that they would advertise because it was seen as industry standard and created "warm feelings" with their customers. So at most, I would get some small time customers, while getting paid much less.
Now I just work on stuff (and yes, I am working on the issue at hand with existing models), and publish it to github (not gonna share it cause don't want my HN account associated with it). If someone contacts me with a dollar figure Im all game.
Nothing shows lack of understanding of the subject matter more than referencing the Dunning Kruger effect in a conversation.
Of course, the smaller models aren't as good at complex reasoning as the bigger ones, but that seems like an inherently-impossible goal; there will always be more powerful programs that can only run in datacenters (as long as our techniques are constrained by compute, I guess).
FWIW, the small models of today are a lot better than anything I thought I'd live to see as of 5 years ago! Gemma3n (which is built to run on phones[2]!) handily beats ChatGPT 3.5 from January 2023 -- rank ~128 vs. rank ~194 on LLMArena[3].
[1] https://blogs.novita.ai/what-are-the-requirements-for-deepse...
[2] https://huggingface.co/google/gemma-3n-E4B-it
[3] https://lmarena.ai/leaderboard/text/overall [1] https://blogs.novita.ai/what-are-the-requirements-for-deepse...
No. They released a distilled version of R1 based on a Qwen 32b model. This is not V3, and it's not remotely close to R1 or V3.2.
We're around 35-40 orders of magnitude from computers now to computronium.
We'll need 10-15 years before handheld devices can run a couple terabytes of ram, 64-128 terabytes of storage, and 80+ TFLOPS. That's enough to run any current state of the art AI at around 50 tokens per second, but in 10 years, we're probably going to have seen lots of improvements, so I'd guess conservatively you're going to be able to see 4-5x performance per parameter, possibly much more, so at that point, you'll have the equivalent of a model with 10T parameters today.
If we just keep scaling and there are no breakthroughs, Moore's law gets us through another century of incredible progress. My default assumption is that there are going to be lots of breakthroughs, and that they're coming faster, and eventually we'll reach a saturation of research and implementation; more, better ideas will be coming out than we can possibly implement over time, so our information processing will have to scale, and it'll create automation and AI development pressures, and things will be unfathomably weird and exotic for individuals with meat brains.
Even so, in only 10 years and steady progress we're going to have fantastical devices at hand. Imagine the enthusiast desktop - could locally host the equivalent of a 100T parameter AI, or run personal training of AI that currently costs frontier labs hundreds of millions in infrastructure and payroll and expertise.
Even without AGI that's a pretty incredible idea. If we do get to AGI (2029 according to Kurzweil) and it's open, then we're going to see truly magical, fantastical things.
What if you had the equivalent of a frontier lab in your pocket? What's that do to the economy?
NVIDIA will be churning out chips like crazy, and we'll start seeing the solar system measured in terms of average cognitive FLOPS per gram, and be well on the way toward system scale computronium matrioshka brains and the like.
Intel struggled for a decade, and folks think that means Moore's law died. But TSMC and Samsung just kept iterating. And hopefully Intel's 18a process will see them back in the game.
I suspect many people conflated Dennard scaling with Moore's law and the demise of Dennard scaling is what contributes to the popular imagination that Moore's law is dead: frequencies of processors have essentially stagnated.
Chiplets and advanced packaging are the latest techniques improving scaling and yield keeping Moore alive. As well as continued innovation in transistor design, light sources, computational inverse lithography, and wafer scale designs like Cerebras.
Of course, feature size (and thus chip size) and cost are intimately related (wafers are a relatively fixed cost). And related as well to production quantity and yield (equipment and labor costs divide across all chips produced). That the whole thing continues scaling is non-obvious, a real insight, and tantamount to a modern miracle. Thanks to the hard work and effort of many talented people.
Wikipedia quotes it as:
> The complexity for minimum component costs has increased at a rate of roughly a factor of two per year. Certainly over the short term this rate can be expected to continue, if not to increase. Over the longer term, the rate of increase is a bit more uncertain, although there is no reason to believe it will not remain nearly constant for at least 10 years.
But I'm fairly sure, if you graph how many transistors you can buy per inflation adjusted dollar, you get a very similar graph.
https://imgur.com/a/UOUGYzZ - had chatgpt whip up an updated chart.
LoAR shows remarkably steady improvement. It's not about space or power efficiency, just ops per $1000, so transistor counts served as a very good proxy for a long time.
There's been sufficiently predictable progress that 80-100 TFLOPS in your pocket by 3035 is probably a solid bet, especially if a fully generative AI OS and platform catches on as a product. The LoAR frontier for compute in 2035 is going to be more advanced than the limits of prosumer/flagship handheld products like phones, so theres a bit of lag and variability.
Not sure about the stated GFlops.. but I suspect we find that AI doesn't need that much compute to begin with.
Well, these days people have the equivalent of a frontier lab from perhaps 40 years ago in their pocket. We can see what that has done to the economy, and try to extrapolate.
The current models are simply inefficient for their capability in how they handle data.
if you base your life on Kurzweil's hard predictions you're going to have a bad time