Also Nvidia isn’t really creating money. The 500B number is third party capital that already exists (BX, Apollo, etc).
Also Nvidia isn’t really creating money. The 500B number is third party capital that already exists (BX, Apollo, etc).
The reason Nvidia is comfortable making these deals is because if OpenAI can’t use the compute, someone else can.
Granted OpenAI going insolvent likely means a drop in the value of compute…
I see two factors converging to cause a collapse of this house of cards:
1. People are realizing that what they need isn't more general intelligence, it's more specialization. A small but well tuned coding model, a small but well tuned customer service model, a small but well tuned document explorer.
2. Specialized hardware - TPUs and NPUs - especially coming out of china. The latest GLM model was trained and runs on Huawei hardware. Nvidia is only worth so much because they are the biggest and best provider of the kind of compute needed to run llms, but the export bans mean china has a lot of incentive to topple that monopoly.
The amount of compute we need to do the things llms do is falling rapidly, the number of people who can provide that compute is rising.
It’s not quite as simple as that. Several studies have shown the opposite: models trained on more diverse knowledge tend to cross-pollinate across domains. So a more generalized model can actually perform better than a specialized one.
That’s why you’re not seeing tons of tiny models (one for Python, one for Pascal, one for Rust, etc).
But it doesn't match my experience. Qwen3.8 27b is clearly smarter at coding than MANY bigger models. gpt-oss-120b for example, is almost 4x the size, and performs way worse at coding tasks.
It's clear to me that you can build small models that work well at specific tasks.
Python vs Rust is probably too fine grained a way to build a model. Coding in general seems like a better target.
There will always be a place for large generalist models, no doubt. But I think that place is much smaller than the big ai companies are counting on.
It practically became a joke about how a huge amount of the training data for GPT-4 was bottom of the barrel reddit vomit and obvious bot spam. Leading to many bizarre edge cases.
Shows qwen3.8-27b along side seven larger models of ~similar vintage. Only one scores above 27b.
Many of those are closed models so idk their exact parameter count / active param count, but it hardly matters - i’m sure all of them are far above 100b params
My point is not that bigger is pointless. It’s just clearly not the only road to take to make a model better, which is obvious just from seeing how models of the same size have gotten better over the past few years
First off, I'd include Qwen flash-next and GLM 5.3 to show some of the other strong open weight models, and they predictably dominate it, but they're much larger. But, it shows up right next to DSv4 Flash 0731 on the overall index, and that's much larger. It's a great model! But then scroll down and hit Time Per Task, and you'll see that DSv4 Flash takes 3.6 seconds per task to Qwen's 21.1. That's what I meant when I said this:
>speed due to excessive thinking maybe to make up for the smaller amount of world knowledge baked in (qwen 27b's main issue iirc), etc - they're tuned for different things.
It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens. And that's a great trade if you don't have the vram to run the bigger models, but speed is pretty important for getting things done... And that's why DSv4Flash is great, too, despite being much larger, and scoring similarly on the intelligence index.
Absolutely - no argument from me here. Bigger is very clearly a lever you can pull to get more out of a model.
> But then scroll down and hit Time Per Task, and you'll see that DSv4 Flash takes 3.6 seconds per task to Qwen's 21.1
Fair point, qwen definitely is slower - it’s a dense model, 27b params, vs a sparse 13b active params model - but the data doesn’t quite agree with what you’re saying about reasoning. I.e.:
> It can make up for its shortcomings by iterating a lot longer, and using way more thinking tokens
If you look at the total tokens generated, deepseek thought for 45k tokens and qwen thought for 48k. Barely a difference. The wall clock difference is all down to the speed of token generation, not the amount of reasoning done. At least when we are comparing deepseek and qwen 27b. The comparison swings more towards your position when it comes to the other models on the chart that reason for much fewer tokens.
So perhaps a hypothetical Qwen-27b-a13b could never rival deepseek’s larger model and the tradeoff is one of speed vs overall size - i.e. a small model needs more active params to compete than a big one does.
One data point that seems relevant to me is that the previous gen qwen Qwen3.6-27b was not so different in performance from its sibling model Qwen3.6-35b-a3b. We never got a qwen3.8-35b-a3b, but if we had, would the gap have stayed the same or gotten bigger? I.e. would the quality gains by improving training coming up against a hard limitation with 35b, or not.
>One data point that seems relevant to me is that the previous gen qwen Qwen3.6-27b was not so different in performance from its sibling model Qwen3.6-35b-a3b. We never got a qwen3.8-35b-a3b, but if we had, would the gap have stayed the same or gotten bigger? I.e. would the quality gains by improving training coming up against a hard limitation with 35b, or not.
Yeah good question, kind of shocking that a 3b active model would perform as well as a 27b dense.
I make heavy use of smaller local models on a daily basis (Qwen3-VL for auto-captioning images, Gemma3:27b for some translation work, etc.). Gemma3:27b is a good example of a very capable general purpose multimodal model and has handled almost everything I've thrown at it from sentiment analysis to documentation writing.
I suppose I was drawing a distinction between specialized and general intelligence versus small and large. I don’t think those are necessarily mutually exclusive.
And Qwen3.8-27b is still better at coding than opus 4.1.
Yes, if you list off models 27b is better than it’s all older models. But that’s my point - newer models are better than older models at the same AND much smaller size. That’s because model size matters less than they say. Training data and model architecture matter more.
In 2026, the default outlook should be suspicion for any big private organisations with profit motive.
You don’t need to think about climate change studies. Instead you can read the allegedly tainted studies we’re actually talking about and profess to all of us what is wrong with them. You can’t point to exactly where they’ve fudged them.
And let’s be real. If this were some conspiracy that OpenAI, Google, and Anthropic were tacitly or overtly conspiring on, I’m pretty sure Elon or Zucc would ruin it to get ahead
Yeah, the cross domain transfer learning from RL is overstated by a lot.
And even if compute demand were perfectly elastic it’s only a good thing insofar as it drives demand for new Nvidia hardware. If tokens can be served from Apple hardware or Google hardware or Huawei hardware that doesn’t help Nvidia.
Across that period, first world countries massively increased their demand for light.
Further efficiency increases only matter to individual choices if they take a use case from {economically impossible} to {economically possible}.
I'd offer that by the 1920s, most goings-on in first world urban environments were no longer price constrained in terms of their light usage.
[0] See table 1.4, p21 https://www.nber.org/system/files/chapters/c6064/c6064.pdf
I mean… some homes definitely do. You must have seen those houses that are all lit up front the outside by lawn mounted spotlights.
If that dies because a lot of people’s needs turn out to be met by a system at home they can run a 30b-150b model on, a lot more of that money goes to apple or intel or amd.
> it's more specialization
China, constrained by hardware, and talent (not to slight the Chinese, but they are limited to domestic resources - and much of the US effort is very international). They did, what the Chinese do, and optimized the process of production, and drastically lowered the cost of development of their models. Cheeper to build, cheaper to run is just good economics.
Meanwhile in the us, we have open AI doing "experiments" - it looks like the costs around the hugging face hack are going to be about the same as China would spend on building out one of their smaller efforts (several million dollars). (Depending on whos numbers you trust, the fact that I can even make this claim should make you raise an eyebrow).
Go back to the 80s' and "expert systems" - most people will tell you that for their time, they were amazing, and useful. People would have loved to have more of them but they were so cost prohibitive that we all but abandoned them for serious use. The US frontier labs seem to have forgotten this lesson and their calls to "slow down" look like an excuse to "cut the waste so we can move to making money".
> Compute has already lost value for me. Six months ago I thought you needed a 1T+ model to be useful coding. Now I am able to get by just fine with a 27b model.
and I was explaining why the complexity question is important in that context. Specifically, because it's part of why people judge the 27b model to be adequate compared to the 1T+ one.
Yes, lots of people are slopping out low-quality things with powerful models because they don't know how software quality works and expect magic. I think we are talking at cross purposes but have no real disagreement.
Just like how I could borrow $100 from you, and you could borrow $100 from me, and we'd be $200 in debt in total, I could buy 10% of shares in your company for $1m and you could buy 10% shares in mine, we'd have 2 companies worth $20m together in total.
1. Yes, smaller models will become more popular, especially as the tokenmaxxing trend dies down and people start stretching their budgets farther. That is a downward pressure on demand.
But along the same dimension, consider that currently only about 40 - 60% of the world uses AI for only about 5 - 15% of their work hours. That means there is still 2x growth from users and 7x - 20x growth from the rest of the work hours left to capture! That is 14x - 40x more demand. Then consider that agentic tasks require multiples more tokens, and that is the kind of usage that is most likely to be deployed, and also the kind of usage that is the least used right now. That's another huge multiple to be tacked on.
And the entire AI industry has been lamenting the extreme compute crunch they're facing (and also why Claude has 9's comparable to GitHub; whereas OpenAI has been chugging along because Altman was OK being called a "podcasting bro" while desperately scrounging for compute years in advance.)
Nvidia's meteoric rise is entirely due to this kind of exploding demand with extremely limited supply.
2. Competing hardware is definitely a threat, but it has its own hurdles. Because the real bottleneck is not Nvidia, it's TSMC.
Pretty much all demand for all chips in all devices in all the world flow to, like, 3 companies in the world that actually fabricate them, and TSMC is the biggest. And the supply is extremely tight, as the exploding costs of electronics clearly shows.
So now TSMC will of course try to keep all its customers happy, but it will inevitably be forced to choose which ones it will keep happiest. And those will be the customers who can pay it the most. And that would be the one with all the money from its de facto status as a monopoly (and possibly even a monopsony)...
Which would be Nvidia ;-)
So yes, compute per task is falling rapidly... but it's barely a dent in the humongous total addressable demand, and the amount of hardware to support that compute is still very constrained, and most of that supply will likely flow through Nvidia.
Ah yes, i am constantly lamenting that my barista isn’t using ai enough ;)
Hopefully you’ve adjusted your ceiling numbers to account for the large amount of people who can’t afford to pay for llms, and will never be able to pay, and aren’t worth it to advertise to since they can afford very little
It's not the baristas or other workers who are spending that money.
(You can explore adoption rates in various industries, including "accommodation and food services", here: https://www.genaiadoptiontracker.com/#explore-data)
If AI makes knowledge workers 1% more efficient on average that is $500 - 700 billion annually. Actual productivity numbers from studies from all the way back in 2024 put the productivity boost above 30%, so add the appropriate grains of salt and adjust numbers accordingly.
That is what the ceiling numbers should be based on. Your baristas will also be using AI eventually, paid for by their employers of course, but that would just be a cherry on top of the real TAM.
The problem with that is that OpenAI can only afford to pay for the compute because they are burning investor money (and so are most of OpenAI's biggest clients). They are losing billions. If they stop burning money, nobody else will be there to pay for that compute at OpenAI's cost.
Sure, somebody will probably be able to use these GPUs, they just won't be able to pay nearly as much for them as OpenAI does.
In reality, it's just nowhere near worth as much as OpenAI pays for it. Inflating the cost of compute is part of the problem caused by the circular financing, and if (or maybe when) OpenAI goes, the price of compute will go with them.
But that’s the point. Investors believe investment in AI will pay off.
So when people say "somebody else will buy that compute", they are correct, but they are ignoring that it will still cause a massive decrease in revenue generated by that compute.
If you're asking me what my point about "believing" is, it's this. OpenAI is losing billions every quarter. It's only staying afloat because it can still find investors who believe in its ability to eventually become incredibly profitable. But there are signs that this is starting to change (e.g. it's unlikely that Softbank will be able to find much more money to give to OpenAI). If OpenAI can't keep raising more money and can't IPO at the level they need (as seems to be the case, given that they keep pushing any IPO date forward), OpenAI will eventually run out of money.
I'm not sure how investors believing in things changes any of this, so now it's your turn to explain what the point is you're actually trying to make.
Failure to take into consideration those kind of correlations ("If my biggest client isn't able to buy it, I would be able to find someone else who will") is one of the principle causes why many risk models turned out to be garbage during the Great Financial Crisis.
But I also doubt Nvidia is on the hook if OpenAI just no longer wants the compute. I bet they are only on the hook if OpenAI cannot pay for it (is insolvent in some way).
I also have to bring up that OpenAI has already spat out an inference chip that beats Nvidia on flops per watt. So they could potentially not need the compute while other ai companies do.
How? The hardware is in OpenAI's datacenters. Does Nvidia have a couple hundred semi trucks, contractors, and IT technicians, to repo the hardware and resell it to someone else before it's lost most of its value? These chips will be replaced approx every 3-4 years. So if OpenAI tanks, after Nvidia pays for and waits for the process to collect the hardware, they then have to sell it for pennies on the dollar. They lose almost all the investment.
Also consider that SpaceXAI already had datacenters full of gear that they basically weren't using because nobody wanted their product, so they now rent it to Anthropic. The demand for hardware isn't really there at the scale of OpenAI.
Via logistics contracts, yes.
All the "frontier" AI companies *are* currently insolvent. They have never been anything other than cash burning machines.
The only way they keep the lights on and the doors open is by borrowing money --- and epic amounts of it. If those operating the cash spigot decide to turn it off, all AI companies will likely be similarly affected --- and so will Nvidia.
OpenAI expects to burn through more cash between 2024 and 2029 than Uber, Tesla, Amazon and Spotify did - combined - before those companies started making money
https://www.morningstar.com/news/marketwatch/20251205243/thi...
Fairly sure data center construction costs are also going up (they require so many resources that everything is constrained at the moment, especially electricity production).
So I don't understand in what world these frontier AI companies can somehow become profitable. The basic tech they're using is basically the same. Yes, around the edges there are a lot of things that can be done, and were done, like caching, batching, mixture of experts, etc, but basically everyone has done all of that by now, and they're still losing money.
So:
Total costs going up a lot - revenues per unit not increasing proportionally, if anything, Chinese models are forcing those down.
How does that math work out to profits? I don't see it.
Or about as bad, after trillions of dollars in investments over multiple years, let's say the entire frontier AI sector has a total profit of $20bn by 2030. In what world does that make sense? Assuming they can scale that total profit to $100bn in 2035 without investing another cent from 2027 to 2035 (utterly ridiculous), the return on investment would happen in roughly 20 years.
China is the one that is really in the driver's seat here. They have the opportunity and the ability to nullify/wipe out our huge investment in AI.
They would have to go insolvent in a way that hits Nvidia revenue. Those are related by distinct factors, a difference that may matter in a crisis.
Gaming is lucrative but not THAT lucrative…
They have no idea how much compute will be needed. They are estimating and trying to push it but they do not know the future.
That’s what loans are.
If NVDA gives a loan, its a swap on the active/left part of their balance sheet - thus, this does not have an effect on overall(!) money suppley. The balance stays the same, also the total amount of money in the whole system at that current point in time.
If a bank gives a loan, then it is extending/enlengthing its balance sheet. (the money created,though,needs to be available on the central bank account to float to other banks when paying for something,therefore we have liquity rules in reg reporting in a bank etc.,meaning someone else needs to have increase thecentral banking liquidity at another part in the system, see things like: https://en.wikipedia.org/wiki/Maturity_transformation )
It's not. It's stating a condition. For traditional banks, a stock crash can independently trigger a failure.
> whole premise of the circular financing worry is that Nvidia sits in the middle of all the guarantees made to companies like OpenAI. If any of those companies become insolvent, Nvidia is on the hook for it
Sorry, I meant revenues. If Nvidia's revenues stay stable, these commitments aren't a problem. Even if the stock price crashes.
> Nvidia isn’t really creating money
It absolutely is. Similar to the way banks create money [1]. The commitments support credit that wouldn't exist without it.
[1] https://www.bankofengland.co.uk/-/media/boe/files/quarterly-...
If NVDA gives a loan, its a swap on the active/left part of their balance sheet - thus, this does not have an effect on overall(!) money suppley. The balance stays the same, also the total amount of money in the whole system at that current point in time.
If a bank gives a loan, then it is extending/enlengthing its balance sheet. (the money created,though,needs to be available on the central bank account to float to other banks when paying for something,therefore we have liquity rules in reg reporting in a bank etc.,meaning someone else needs to have increase thecentral banking liquidity at another part in the system, see things like: https://en.wikipedia.org/wiki/Maturity_transformation ) ---
As working in the field, I know the BOE paper very well ;-)
But aren‘t they currently sitting on like a margin of over 80%? So I assume their revenue will fall, because people are not going to like that in the long-term.
Making a loan/offering credit isn't automatically money creation - the amount of money in the system before and after the loan might be the same. Haven't been following Nvidia all that closely, but it seems a little bit unlikely that they're a commercial bank. Financial chicanery they may be doing but offering deposit accounts would be new territory. The loan has to be made in a very particular way for it to be money creation (notably, in a way that creates new money), and it should be illegal for most people to do that otherwise we'd all be printing our own money instead of the printing being directed to wealthy asset owners first and foremost.
The Fed can limitlessly increase monetary base (MB) to increase M1. Non-bank lenders create M4 which turns into M2 through the money markets which turns into M1 through the banks. More convoluted and more limited, by design. But the net effect is the same–more M1.
More broadly: when you open a bar tab, you're literally creating money, even if it never hits M1.
Credit is always money creation. When it's done at scale it's almost always creating M1.
No, its not! Its a swap on the active/left side of the balance sheet: If my mum borrows me 1 USD for icecream, she did not create money, instead she swapped her positions in her assets/left side of her balance sheet.
Or I suppose I should spend more time with an open bar tab come tax time.
The monetary aggregates are things that are theoretically equivalent to creating money, but I'm going to challenge whether it is proper to compare M4 activity to the fed balance sheet. One of those entities is involved in actual money creation. One of them is hyper-privileged money creation that is more impactful and less comparable to other forms.
Yes.
Nvidia issues a commitment to a firm that lets an SPV raise money. Some of that is just M1 being transferred from one account to another. A lot, however, will be created through loans (new M1), commercial paper (new M3) and the like. Some of that will be pledged as collateral. On the other side of the equation, when the SPV signs e.g. a construction contract, you'll have builders taking out bank loans and issuing commercial paper, et cetera.
You can pay your taxes out of a checking account using money a bank created out of thin air that morning that the Fed won't even learn about until the evening.
> monetary aggregates are things that are theoretically equivalent to creating money, but I'm going to challenge whether it is proper to compare M4 activity to the fed balance sheet
You're absolutely correct in M4 being unequal to the monetary base and deposits at the Federal Reserve. But the moment you say the Fed balance sheet, you're conflating two things. The Fed has lent, in the past, against non-Treasury collateral (most famoulsy, mortgage bonds).
And if we're talking about money in relation to the real economy, you can't spend reserves at the Fed. There is debate on what the most "real" money is, but pretty much everyone agrees that a deposit in a checking account is more relevant to the economy than unspendable reserves at the Fed.
How? And what is your theory why people pay taxes from their labour earnings instead of just creating new money to pay Just-In-Time?
> A lot, however, will be created through loans (new M1)
Isn't M1 is currency and overnight deposits? I'm pretty sure you're just making things up here; it sounds highly illegal for Nvidia to create new M1. Unless they've taken out a banking license when I wasn't looking.
> You can pay your taxes out of a checking account using money a bank created out of thin air that morning that the Fed won't even learn about until the evening.
Ditto prior. Nvidia isn't allowed to do that.
...from a checking account. Banks thoughtlessly lend against investment-grade commercial paper. The kind the SPVs Nvidia is commiting to are borrowing with. Nvidia's promises are creating debt that money markets and banks transmit into M1.
Note that the Fed doesn't directly create M1. It moderates it through the interbank lending market to manipulate the monetary base, M0 plus deposits at the Fed. That, in turn, influences banks' lending patterns which is what manipulates M1 and MZM. Underwriting M4 to drive up collateral that in turn increases M1 is the same mechanism, different channel.
> what is your theory why people pay taxes from their labour earnings instead of just creating new money to pay Just-In-Time?
I'm not Nvidia. I did just sign a purchase order with a builder for my deck that their local bank called me to confirm before issuing a loan that will appear as new money in their checking account. My signature, in that case, enabled the bank to create a tiny amount of money. It's not monetarily significant, however, because my deck isn't that fancy.
At the end of the day, all credit is money. If you can create credit, you can create money. Most of us can't, at least not in significant quantities.
> Isn't M1 is currency and overnight deposits?
M0 + demand deposits at commercial banks, yes.
> sounds highly illegal for Nvidia to create new M1
Nvidia causes the creation of new M1 through banks. Nvidia's commitments directly create money of a quantifiable amount that wouldn't have existed if Huang hadn't flicked his pen.
> unless they've taken out a banking license
Fun fact, you don't need a banking license to issue traveler's checks. And traveler's checks are counted in M1 in the U.S.
> pretty sure you're just making things up here
If that's your reaction to encountering new information, godspeed.
We've got Schrodinger's money here! One minute I can go down to the bar and create as much money as I like with the bartender. Then the next we suddenly discover that yes we need a bank and yes there all of a sudden a checking account is involved and no a bar tab isn't a legally recognised form of money.
> I'm not Nvidia. I did just sign a purchase order with a builder...
Again though, we discover that in fact the person in control of the part where the actual money was created was a bank. You can't just head off with your builder and create money - otherwise you'd be stupid to stop until you unseat whoever was the richest fellow in the world this morning.
> Nvidia causes the creation of new M1 through banks.
And the banks make another appearance!
---
I put it to you that Nvidia can't, in fact, create new money. You'd never get that line of argument past a tax collector.
> Fun fact, you don't need a banking license to issue traveler's checks. And traveler's checks are counted in M1 in the U.S.
And I'll add in postcript that I know nothing about travellers checks in the US, but I'm quietly confident it'd turn out to be illegal to go around creating vast amounts of new money there too, on the basis that people tend to work for a living.
Load bearing, heavy lifting... Your comment wasn't LLM-written, either. I think we're starting to see LLMisms infect human writing. I might try to start speaking like this and see if anyone notices. It could be a good gag.