> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Wow!
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Wow!
It'll all be as cheap as DeepSeek was before the price hike. And, it'll become more and more realistic to run near-frontier intelligence on personal devices.
I'm honestly baffled they were not acquired by somebody else (sorry AMD).
The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years.
Cerebras will win in terms of approach.
It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.
Yes, please!
Slightly disagree. It really depends on the price-point at which they can do that etching. ~1k usd / ~30B model in a hdd-sized case that fits on your desk? I'd buy one right now, even knowing that I'm "stuck" with whatever model of the day is.
> From the moment a previously unseen model is received, it can be realized in hardware in only two months ( https://taalas.com/the-path-to-ubiquitous-ai/ )
Qwwen3.5 122b was released 6 months ago and is still best in class overall 100-140 B param model.
....
:T
Considering the 8B model uses 53 billion transistors, that's 6.625 transistors per parameter.
Assuming they can get it down to 3 (somehow), that's still 300 transistors, or 5.565 RX 9070s.
https://www.techpowerup.com/gpu-specs/radeon-rx-9070.c4250
You're looking at
1) waiting for another 3-5 generations of transistor improvements before it can fit into a single conventional chip, or
2) another generation before getting a monster of a chip (1000+ mm^2), and prices for flawless etching scale quadraticly (likely $1000+ for manufacturing costs alone).
Could happen, but it's a long shot for a market that could be satiated by specialized accelerators.
What costs are you talking about and why would them be a problem?
If it is the price: «Kharya says it costs 100x as much to train a model then to get a customize HC chip in reasonable volumes from Taalas» ( https://www.nextplatform.com/compute/2026/02/19/taalas-etche... ).
That is thinking about an LLM (logical) producer and server. For mass production, the costs go down. And in the case of a ~100b model as the poster mentioned, they would be just single PCI cards with 5 or 6 HC2 chips: doable and practical.
I see more potential problems in the positioning of the SRAM - but not a real problem given that excellent team.
To get a proper idea of the costs the architecture of the HC2 will have to be clearer.
There's a bit of a ticking time bomb there, something that a Taalas-like architecture can clearly resolve.
I distinctly remember 32-bit/33 MHz PCI accelerator cards for SSL being a real thing (for use on OpenBSD or FreeBSD), in an era when something like a single core 700 MHz Pentium 3 1U system was a relatively powerful individual bare metal httpd box.
http://www.aster.si/partnerji/compaq/atalla/axl200.html
The CPU load of doing a lot of SSL purely in software was a problem in terms of scaling things up, so this was one attempt at a (very short lived) solution. Note that this predated TLS1.0.
We can have that discussion now: sounds like that would kill OpenAI and Anthropic
Stack on top of that the fact that diffusion based models like the ones made by Inception Labs are far faster and more efficient than autoregressive LLMs and have an even higher ceiling of optimization (single step path prediction via model distillation versus 50 step denoise is currently an active area for image diffusion)
The human brain is soon neither going to be more powerful nor energy efficient than the stuff we use to run AI.
GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.
In that 5+ year timeline, the compute per watt could change by three orders of magnitude.
GPUs are to LLMs what CPUs are to gaming — not a good fit.
If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, different ALU design, model architecture changes, etc.
If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.
But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.
In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system?
And maybe llm's could invent. So even more to simulate. And simulation is inherently compute-heavy.
So unless there are some other bottlenecks, we'll use a lot of simulation servers.
A rough analogy would be if the first generation of ISP's spent billions on dial-up exchanges, when fibre could be invented next year.
He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]
The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.
New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.
New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.
In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.
We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.
2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation.
> If Amazon doesn't buy but Microsoft does
The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now.
I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.
2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens.
3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC.
Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.
2. Jevon’s paradox is about total consumption, not margins. Valuations are about margins (and their projections). Many coal mine owners went bust despite increased total coal consumption.
3. Source? Gemini for example is 70% on TPU. I have yet to see data on Maia-300 beyond Microsoft PR. Remember it doesn’t have to be better it has to be more cost efficient. The overwhelming majority of inference spend does not care if token output is 20% slower if it is 50% cheaper.
> Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.
What Nvidia needs. Whether neoclouds can stay competitive vs hyperscalers paying Nvidia tax is far from clear, particularly when inference margins compress.
2. Total consumption drives more demand for the already supply constrained hardware. Can AI hardware market go bust? Sure it can. But being early is the same as being wrong in the investment market. When do you predict the bust to be?
3. Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can. The biggest tell on how Nvidia is doing is that their share in inference has increased despite the increase in competition: https://archive.md/CKP0N. So while competition is getting bigger and bigger because the overall pie is getting exponentially bigger, Nvidia's growth is still higher than average.
Some other sources:
https://www.businessinsider.com/amazon-nvidia-aws-ai-chip-do...
https://www.businessinsider.com/startups-amazon-ai-chips-les...
Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.
> Total consumption drives more demand for the already supply constrained hardware.
Demand is the wrong metric.
Only number that matters is whether AI attributable revenue will be sufficient to pay back enough of each successive capex vintage (e.g. 750B this year, 1T next year, 1.2T in 2028) so that hyperscalers and neoclouds can either self-fund or continue to issue debt as bond markets are already straining and tax-payer backed sovereign debt is providing a high baseline. Otherwise they downgrade capex projections and the bubble pops.
Expensive compute needs expensive inference to justify 30-40B/year/GW of compute. There are many reasons why frontier API pricing which is what the industry is based on may not persist. It is also almost certainly the case that 2026 is the worst year of supply and demand mismatch to allow for 80%+ margins. HBF next year has the potential to single handedly pop the DRAM spot bubble.
> Can AI hardware market go bust? Sure it can.
This is the bear thesis. It is not that AI will crash or be useless.
> But being early is the same as being wrong in the investment market. When do you predict the bust to be?
Q4 27-Q2 28 is when the bill becomes due at the latest. There are sufficient financial levers left to buy time without returns until then.
> Google, Amazon, Microsoft, Meta are all buying as many Nvidia GPUs as they possibly can.
All of these companies have rock solid revenue streams and can easily swallow 500B of capex devaluation over time. Their buying of Nvidia today is not necessarily the indicator you are implying as there are strong competitive reasons to make the game more expensive for everyone else.
Burry’s main argument is depreciation is being understated and the capex vintages will not be paid off before they are essentially useless. This can happen whether or not aux is reused.
And why does he think depreciation is understated? It is because he thinks newer Nvidia GPUs will make older ones obsolete faster. Hence, my entire post.The rest of your argument centers around whether AI growth will meet the cap ex expenses. I don't see anything new in it.
HBF next year has the potential to single handedly pop the DRAM spot bubble.
I'll believe it when I see it. Jevons paradox will apply here again in my opinion. HBF does not replace HBM.Why? Make your case.
There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but more needs to be discovered so that we can reduce building hundreds of more data centers as the alternatives mature.
As better software becomes more useful for the alternative AI hardware for developers with LLMs running efficiently you then would have more choices of hardware to run your LLMs on rather than just only GPUs.
Unlimited.
What has been the limit to electricity demand globally?
Unlimited.
We can't get enough and never will. Costs have to become pretty severe to turn back the demand as well.
https://openrouter.ai/rankings#top-models
And their market share sits at around 16-20%.
At 75.3 trillion tokens for the week ending 10 Aug 2026, that means that up to 450 trillion tokens were plausibly demanded by the whole market for that week.
My take: At max saturation, each person on earth could have their demands satiated by an average of 16 agents running concurrently. Sometimes more, often times less, but the average would likely be at 16.
At 200 tokens/second for each agent, that would mean 15.48288 quintillion tokens per week.
We're currently at about 0.00290643601% of the calculated demand ceiling.
Even if the demand limit per person is just 1 agent at 50 tokens/second, the current demand's still 0.186011905% of the theoretical ceiling.
They're typically not built where you want housing, and the buildings are distinctly the wrong shape.
If you can't use the power infrastructure profitably my next thought would be warehousing.
But also... we've seen a pretty continually increasing demand for compute. Even if AI busts a bit (or becomes a bit more efficient) I bet most data centres stay data centres, just less profitable ones.
> built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling.
Insurmountable according to whom? And who assume that only GPUs are all we need to continue scaling? Google, Amazon, Microsoft, Meta and OpenAI, all have or plan custom non-GPU AI chips. Do they plan to use them not for scaling?