Nvidia CEO Jensen Huang announces new AI chips: ‘We need bigger GPUs’
cnbc.com
cnbc.com
Obviously they are going to keep doing bigger. But the takeaway for me is that they are building "docker for llms" - NIM. They are building a container system where you can download/buy(?) NIMs and easily deploy them on their hardware. Going to be fun to watch what this does to all the AI startups...
What specific class of AI startups do you have in mind here? AI-aaS startups who provide the "infra"?
Generally if whatever AI product you have can easily just be a feature in whatever application businesses already use, then you are running a business on borrowed time.
Proper background removal of even remotely complex content is still in demand, especially on a large scale. I doubt you'd use an iPhone to work on >100 of images per second.
https://findthatmeme.com/blog/2023/01/08/image-stacks-and-ip...
PhotoRoom begs to disagree with this statement.
I can understand from a growth perspective why it’s more profitable for Nvidia if it can become more of a platform service for AI. However, that’s difficult to balance that and partnerships the company already has with AWS and Microsoft. I’d expect to see some acquisitions or competing custom solutions in the future. Fortunately for Nvidia, a lot of AI is still dependent on CUDA. I’m interested to see how this plays out.
NVIDIA could voluntarily open the standard to avoid this ligitation if they wanted to, though, and IMO it would be the smart thing to do, but almost every corporation in history has chosen the ligitation instead.
The MI300 smokes the H100 yet here we are.
CUDA is such a misnomer. Amd doesn't have tensorRT, cuDNN, cutlass, etc. Forcing Nvidia to make these work on AMD is like forcing Microsoft to make windows work on apple hardware... Not going to happen.
You can absolutely use the names of the functions and the programming model. Like I said, HIP is literally a copy. Llama.cpp changes to HIP with a #define, because llama.cpp has its own set of custom kernels.
And this is what I've said before, CUDA is hardly a moat. The API is well-known and already implemented by AMD. It's all the surrounding work: the thousands of custom (really fast!) kernels. The ease-of-use of the SDKs. The 'pre-built libraries for every use case'. You can claim that CUDA should be made open-source for competition, but all those libraries and supporting SDKs represent real work done by real engineers, not just designing a platform, but making the platform work. I don't see why NVIDIA should be compelled to give those away anymore than Microsoft should be compelled to support device driver development on linux.
If AMD isn't a competitor before government intervention, I don't the government forcing nvidia to open up CUDA changes much. CUDA's moat isn't due to some secret sauce - nvidia put in the developer hours; and if AMDs CUDA implementation is still broken, people will continue to buy nvidia.
There has been a lot of trying to get AMD to work - Hotz has been trying for a while now[1] and has been uncovering a ton of bugs in AMD drivers. To AMD's credit, they have been fixed, but it does give you a sense of how far behind they are in regards to their own software. Now imagine them trying to implement a competitor's spec?
[1] https://twitter.com/__tinygrad__/status/1765085827946942923
The issue that AMD has is they had a long period where they clearly had no idea what they were doing. You could tell just from looking at websites, CUDA pretty much immediately gets to "here is a library for FFT", "here is a library for sparse matricies". AMD would explain that ROCM is an abbreviation of the ROCm Software platform or something unspeakably stupid. And that your graphics card wasn't supported.
That changed a few months ago; so it looks like they have put some competent PMs in the chair now or something. But it'll take months for the flow on effects to reach the market. They have to figure out what the problems are which takes months to do properly; then fix the software (1-3 months more minimum); then get it into the open and the foundational libraries like PyTorch pick it up (might take another year). You can speed that up, but more cooks in the kitchen is not the way. Bandwidth use needs to be optimised.
It isn't like ROCm seems lacks key features; it can technically do inference and training. My card crashes regularly though (might be a VRAM issue) so it is useless in practice. AMD can check boxes but the software doesn't really work and grappling with that organisationally is hard. Unless you have the right people in the right places, which AMD didn't have up to at least mid 2023.
Pytorch has been "supporting" rocm for all last 2 years
It makes perfect sense that, organisationally, they were focused on that battle. If you remember the Athlon days, AMD beat Intel before, but briefly. It didn't last. This time it looks like they beat Intel and have had the focus to stay. Intel will come back and beat them some cycles, but there is no collapse on the horizon.
So it makes sense that they started looking at nVidia in the last year or so. Of course nVidia has amassed an obscene war chest in the meantime...
If even that. A few years ago they managed to break basic machine learning code on the few commonly-used consumer GPUs that were officially supported at the time, and it was only after several months of more or less radio silence on the bug report and several releases that they declared those GPUs were no longer officially supported and they'd be closing the bug report: https://github.com/ROCm/ROCm/issues/1265
If they paid everyone $1M/year salaries maybe more people would consider going into aerospace engineering.
Right now though Boeing's starting salaries aren't that much higher than what an Uber driver in the bay area makes.
If it was a clean room implementation of the API NVIDIA wouldn’t care. Heck that’s exactly what AMD did with HIP.
But what you cannot do is essentially intercept calls to and reverse engineer NVIDIA binaries in real time because you can’t be arsed to build your own.
And this is precisely what anti-trust ligitation would allow them to do.
Preventing someone from reverse-engineering a product with the sole intention of maintaining monopoly status may be seen as anti-competitive.
AMD makes really, really good CPUs now, but only after ligitation against Intel allowed them to keep up with evolving x86 standards.
It's not about being "arsed" to build your own, the problem is NVIDIA controls the ecosystem-wide standard. NVIDIA can add to CUDA at any point in time and launch a GPU at the same time, the ecosystem would be forced to buy it if they want to stay on the cutting edge, and AMD would never be able to compete or reverse engineer these new standards in time.
What ZULDA did wasn’t to maintain a compatibility with the CUDA API and provided an open implementation of it but rather use all the CUDA based libraries that NVIDIA provided on top of it.
The equivalence again would be that not only Google implemented their own Java compatible API but rather that they used the now Oracle owned JVM to do so and redistributed it.
AMD already implemented CUDA essentially one to one in the form of HIP in ROCm. The issue that they face is that they don't have all the equivalent middleware to make copy pasting code actually work and this was what ZULDA did but instead of building a HIPdnn they just reused NVIDIA binaries.
The problem currently, as people like Hotz and many others are discovering, it not the lack of CUDA. Most people use PyTorch and don't care what the underlying software is. Infact most CUDA is hand tuned to nvidia hardware anyways and is optimized to make the most on nvidia. The problem is AMD's drivers - the piece that actually sends the code to run on the GPU, tends to be broken. AMD cannot "sponsor" an outsider to fix this. A legal, but broken, AMDCUDA will not be any better than the current situation; so no, having CUDA on AMD wouldn't change anything.
The problem is not "CUDA is not AMD", the problem is AMD has not, does not, and for some reason will not invest adequately in GPU compute. CUDA is a mirage; if AMD had a similar platform someone would have done the work already to ensure PyTorch works on it. PyTorch already supports ROCm, people don't use it because the performance is bad and it's buggy. When nvidia had this problem, nvidia hired engineers to work on open source projects and debug issues in open source libraries (not even limited to AI, you will find nvidia engineers debugging issues in a wide range of CUDA projects). When AMD has this issue, they barely acknowledge it.
That’s what Nvidia faces. Doesn’t matter how good the current in-house teams are using direct hardware, the trend in corporate is shift to a vendor (Google/AWS).
Nvidia can watch this inevitable shift or get ready to offer itself as a platform too.
AI graphics service that renders games in the cloud (photorealistic), concluding its epic journey of being an amazing graphics card company.
Kinda cool when you think about it.
It feels like reading that setting up something like Facebook would be extremely challenging for a company like SpaceX.
Just because an organisation is extremely good at one thing doesn't mean it can easily apply that to another field. I would guess that SpaceX probably does have the talent on hand to throw together a Facebook clone, but equally I think they would struggle to actually complete with Facebook as a business at scale.
To give an example in a closer domain: Look at how long Google lost money on cloud services through 2022 (over $15B in loses), and now only makes money by creative accounting (bundling "cloud services" together versus breaking out GCP; Microsoft does something similar with Office 365 and Azure).
Like many potential customers, I would not consider GCP because:
1) Google "support" is a buggy, automated algorithm which randomly thwacks customers on the head
2) Google randomly discontinues products
3) I've seen a half-dozen to a dozen instances where buying from Google was penny-wise and pound-foolish, and so have many other engineers I've worked with.
Google's overall attitude is that I'm a statistic defined by my value to Google. Google can and will externalize costs onto me. That attitude is 100% right for adwords and search, which are defined by margins, but not for something like GCP. If I am going with a cloud service/platform, I'll go with Amazon, Microsoft, or just about anyone else, for that matter.
That's not that Google is a bad company. Google actually did have the skill set to build the software and data centers for a very, very good cloud provider. It's just that Google's core competencies lie very far from providing reliable service to customers, customer support, or all the things which go into providing me with stability and business continuity.
"Fixing" this would require a wholesale culture, value, and attitude change, and developing a core competency very far from what Google is good at.
I put "fixing" in quotes since if you develop too many core competencies, you usually stop being good at any of them. Focus is important, and there's a reason many businesses spin out units outside of their domains of focus. If Google is able to become good at this, but in the process loses their edge in their current core competencies, that's probably a bad deal.
FWIW: I haven't yet formed an opinion on NVidia's cloud strategy. However, their core competencies appear to be very much in the "hard" domains like silicon, digital design, machine learning rather than "soft" ones. Another relevant example for what can happen when hard skills are de-emphasized at engineering-driven companies is Boeing (if you've been following recent stories; if not, watch a documentary).
What I've heard is Azure is a pain in the ass, and things take three times as long to set up there, for some reason. There's also Oracle cloud but you hear way less about them. out there.
> What I've heard is Azure is a pain in the ass, and things take three times as long to set up there, for some reason.
Doesn't matter. The cost here is a rounding error. What does matter is something like this:
https://developer.chrome.com/docs/extensions/develop/migrate...
https://workspaceupdates.googleblog.com/2021/05/Google-Docs-...
https://workspace.google.com/blog/product-announcements/elev...
https://www.tomsguide.com/news/g-suite-free-shutdown
Etc.
These sorts of behaviors take out whole swaths of businesses wholesale. It's random, and you never know when it will happen to you.
It's the difference between managing a classroom with:
- an annoying kid throwing spitballs every day (Azure)
- the quiet kid who, one day, brings an assault rifle, a few extra mags, and starts spraying bullets into the cafeteria (Google).
Yes, one is a constant source of annoyance, but really, it's very manageable when you consider the alternative.
(Oracle, in the school analogy, is the mean kid who spreads false rumors about you. As far as I can tell, there is never a sound, long-term business reason to pick Oracle. Most of the reason Oracle is chosen is they're very good at setting up offerings which align to misaligned incentives; they're very often the right choice for maximizing some quarterly or annual objective so someone gets their bonus. In return, the firm is usually completely milked by Oracle a few years down the line. By that point, the decision maker has typically collected their bonus, moved on, and it's no longer their problem.)
One is a publicly listed business with as much of an objective look at real time "worth" as possible in today's world, and the other is a private business with confidential financials.
Seems like you would be unable to even calculate SpaceX's net worth, much less compare them to a business with the most objective measure of "worth".
A private investment at this scale should have a lot more transparency and due diligence than disclosures from a SEC disclosures. If I were investing $750M, I'd have engineers under NDA review SpaceX technologies, financial auditors, legal auditors, etc.
Secondary sales place it a little bit higher (but those typically have all the issues you describe).
But the fattest profit margins are in selling to big corporations. Big corporations who already have accounts with the likes of AWS. They already have billing set up, and AWS provides every cloud service under the sun. Container registry? Logging database? Secret management? Private networks? Complicated-ass role management? Single-sign-on integration? SOC2/PCI/HIPAA compliance? A costs explorer with a full API? Everything a growing bureaucracy could need. Getting your GPU VMs from your existing cloud provider is the path of least resistance.
The smaller providers often compete by having lower prices - but competing on cost isn't generally a route to fat profit margins. And will folks at big corporations care that you're 30% cheaper, when they're not spending their own money?
nvidia could definitely launch a focused cloud product, that competes on price - but would they be happy doing that? If they want to get into the business of offering everything from SAML logon integration to a managed labelling workforce with folks fluent in 8 languages - that could be a great deal of work.
NVIDIA is not entering an empty marketplace here. Also there's already enough know-how to make things run on cloud, OpenStack, HPC, etc.
Unless they make things very difficult for independent platforms, they can't force their way in that much, from my perspective.
Now they do all re-sell Nvidia GPUs in their cloud businesses, but there the cost will be passed very directly on to infrastructure customers, who will see the competing higher level services from those cloud providers (hosted models) likely at lower or competitive prices, and it's going to be harder to justify renting CUDA cores for custom software.
Microsoft has partnetship with OpenAI and also with Mistral.
Present convenience may not hold true in future. Nvidia knows that well.
Even if AWS has its own hardware+software solution for neural networks it would still take years if not decades to tear off the CUDA platform
Only partially, because in LLMs FP4 isn't half as useful as FP8. So if you have gear that crushes at FP4 then that's what you use and you benefit from that increased speed (at minimal accuracy loss).
Definitely some marketing creativity in there, but its not entirely wrong as a measure of real world usage
...assuming the recent 1.58b paper doesn't render the entire float quantization approach obsolete by then.
- Research for a while now has been finding that smaller weights are surprisingly effective. It’s kind of a counterintuitive result, but one way to think about it is there are billions of weights working together. So taken as a whole you still have a large amount of information.
Wasn't there a paper from Microsoft two weeks ago or so where they trained on log₂(3) bits?
This makes network minimise loss not only with regard to expected outcome but also minimises loss resulting from quantisation. With big networks their "knowledge" is encoded in relationships between weights, not in their absolute values so lower precision work well as long as network is big enough.
4 bits is effectively 16 different float point numbers - 8 positive, 8 negative, no zero and no NaN/inf. 1 bit for sign and 3 bits for exponent, 0 bits for mantissa, mantissa is implied to be 4. It’s logarithmic - representing numbers in the range from -4^3 to 4^3, smallest numbers are 4^-3.
I like how you specified that it's not floating point.
Apparently some people are drawing connections to this paper [1] on 4 bit LLMs, which has one NVIDIA employee among its contributors
Discussed in a previous post https://news.ycombinator.com/item?id=37930663
Ubuntu is actually a pretty great daily driver desktop Linux, and I'd hate for that to lose priority and disappear.
I'm not a fan of what happened to the Red Hat ecosystem for exactly the same reasons.
Probably depends on the team and the worst ones have a lot of churn.
That all said, I have to work with a lot of CentOS and Rocky workstations for VFX and I enjoy those for desktop Linux even less than Ubuntu.
If they defaulted back to a menu and taskbar-based WM, it might actually be more approachable to users who are more familiar with macOS and Windows.
[1] https://www.reddit.com/r/recruitinghell/comments/1bec2zk/lit...
At least I would finally get to buy MS Ubuntu PCs at the shopping mall.
"The computing power needed to replicate the human brain’s relevant activities has been estimated by various authors, with answers ranging from 10^12 to 10^28 FLOPS."
Petaflop is 10^15
Crazy times.
Obligatory Rick and Morty: https://www.youtube.com/watch?v=xerLPWdyX-M
Seems logical since it's becoming impractical to cram so many transistors on a single die.
imagine aws if they also sold all computers in the world, now you can only rent from them
--Nvidia, pretty soon
And that's once you've bought an expensive Tesla/Quadro GPU too.
no sympathy for them.
Hardware: * The GPU * The GPU-GPU Fabric (NVLINK) * The CPU * The NIC * The Network Fabric (infiniband) * The Switch
And that's not even starting to get into the many layers of the software stack (CUDA, Riva, Megatron, Omniverse) that they're contributing and working to get folks to build on.
It is already proven that good language models are possible given enough resources. The challenge now is to put these models in a solution which do not require unfathomable amounts of resources for the average use cases.
This is not a problem with AI only, but with every software we use. Only two groups try to optimize things and try to fit into smaller systems. Passionate programmers and people who is paid to do this (e.g.: phone manufacturers' software teams, etc.).
To your point enthusiasts and developers with strong financial motivation have fairly optimized code. AI definitely falls into the later group. These are not like your typical web app :)
On the other areas where NNs are used, what I can see is training models with less data is not only plausible, but very possible, esp. in image processing, probably an image carries more information than a single sentence.
I think AI falls on the spectrum with a slight bias to the latter group. Because if you can shrink a model 10% and lose a month, you'd rather have a 10% bigger model now, and reap the money^H^H^H^H^H fame.
I don't know anything about a typical web app, because I'm not a typical developer developing web apps. :)
This is very different than a program that requires zero research breakthroughs to dramatically improve and is simply slow and bloated because people have different priorities.
Nope, because all of them are harder than just going bigger.
> Top model performance is an arms race and it's happening at every expense.
This is also what I said. "Growth (in model performance) is king, so quick and dirty (going bigger) beating harder optimization efforts".
> This is very different than a program that... [Snipped for brevity]
Again this is what I said by "it keeps development momentum". Yes people have different priorities. Mostly money and fame in this point.
So, we don't disagree a bit.
I think it is more of an information problem. How can we store enough information in weights so that it is possible to train models without a budget similar to OpenAI
I work on material simulations. I make processors hit their TDPs, saturate their pipelines and make them go as fast as they can. However, sometimes we come up with a formula optimization which does things 1-2% faster, which means we can save hours on a bigger computation. Utilization doesn't change, but speed does.
> I think it is more of an information problem.
It's an interesting point of view, and partially true. However, we're still wasting "space" by just adding bits to the network to make it contain more data.
There's a long way to go. "Wasteful development" is a phase and always be part of software development. The important part is not forgetting that optimization exists. Otherwise we can't sustain ourselves much with all that energy use.
whole conference has been proceeding like this now. look if you invite cramer and the wall street crowd, you should throw in some dollar figures. like - who is paying for all this. how much. and why. talk is entirely about token generation bandwidth, exaflops and petaflops, data parallel vs tensor parallel vs pipeline parallel - do you honestly think cramer knows the difference between an ml pipeline and an oil pipeline ?
i am watching this conf with my kid - proper GenZ member - who got up after 5 mins and said man who is this comedian, his jokes are so bad, and left :(
That’s fine, it’s a developer conference for a founder lead company that hasn’t reached the “stock price is the product” state. He’s not trying to optimize the next 5 days of stock.
There’s a full ecosystem grab with Nim there, a new GPU that forces every major datacenter to adopt (or their competitors will massively increase their compute density)
It's not a hipster event.
And the pauses are clearly his style of presenting. Having a artificial pause to tell the audience it would be a good time to react.
You clearly don't like it, I don't mind it.
But just to be clear: we see peak human ingenuity. This right now will be history on how we as humans build AGI, full human robots etc (in case we don't nuke ourselves).
You can read the conference results easily at any it news portal.
There is no requirement from Nvidia entertaining you or your kid.
That being said, their stock is absolutely and hilariously overvalued.
Apple is in talks with Google to bring Gemini to the iPhone, and it will obviously also be on android phones. So almost every phone on earth is poised to be using Gemini in the near future, and Gemini runs entirely on Google's own custom hardware (which is at parity or better than nVidia's offerings anyway).
Making a graphics chip that is as good as Nvidia: Very difficult. Huge moat, huge effort, lots of barriers, lots of APIs, lot of experience, lots of decades of experience to overcome.
Making something that can run a NN: Much, much easier. I'd guess, start-up level feasible. The math is much simpler. There's a lot of it, but my biggest concern would be less about pulling it off and more around whether my custom hardware is still the correct custom hardware by the time it is released. You'd think you could even eke out a bit of a performance advantage in not having all the other graphics stuff around. LLMs in their current state are characterized by vast swathes of input data and unbelievably repetitive number crunching, not complicated silicon architectures and decades-refined algorithms. (I mean, the algorithms are decades refined, but they're still simple as programs go.)
I understand nVidia's graphics moat. I do not understand the moat implied by their stock valuation, that as you say, they are the only people who will ever be able to build AI hardware. That doesn't seem remotely true.
So... correct me Internet. Explain why nVidia has persistent advantages in the specific field of neural nets that can not be overcome. I'm seriously listening, because I'm curious; this is a deliberate Cunningham's Law invocation, not me speaking from authority.
After 10 years of pretending to care about compute, AMD has filled the industry with burned-once experts who, when weighing nvidia against competitors, instinctively include "likely boondoggle" against every competitor's quote because they've seen it happen, possibly several times. Combine this with nvidia's deep experience and and huge rich-get-richer R&D budget keeping them always one or two architecture and software steps ahead, like it did in graphics, and their rich-get-richer TSMC budget buying them a step ahead in hardware, and you have a scenario where it continues makes sense to pay the green tax for the next generation or three. Red/blue/other rebels get zinged and join team "just pay the green tax." NV continues to dominate. Competitors go green with envy, as was fortold.
More like burned 2x / 3x / 4x of this time it's different people.
Looking at you Intel
But (as a reply to some other repliers as well), AMD was also chasing them on the entire graphics stack as well as compute. That is trying to cross the moat. Even reimplementing CUDA as a whole is trying to cross a moat, even a smaller one.
But just implementing a chip that does AI, as it stands today, full stop, seems like it would be a lot easier. There's a lot of people doing it and I can't imagine they're all going to fail. I would consider by far the more likely scenario to be that the AI research community finds something other than neural nets to run on and thus the latest hotness becomes something other than a neural net and the chips become much less relevant or irrelevant.
And with the valuation of nVidia basically being based not on their graphics, or CUDA, but specifically just on this one feeding frenzy of LLM-based AI, it seems to me there's a lot of people with the motivation to produce a chip that can do this.
I don't think they have a crazy advantage HW wise. Couple of start-ups are able to achieve this. If SW infrastracture end is standardized, we will have a more level playground.
Without CUDA you have a chip that runs on premise without anyone having a clue how good that is which is supposedly what Google does. Your only offering is cloud services. As big as this is, corporations would want to build their own datacenters.
I think nobody had the time to port any of these architectures away from CUDA because: * the leaders want to maintain their lead and everyone needs to catch up asap so no time to waste, * and progress was _super_ fast so doubly no time to waste, * there was/is plenty of money that buys some perceived value in maintaining the lead or catching up.
But imo: 1. progress has slowed a bit, maybe there's time to explore alternatives, 2. nvidia GPUs are pretty hard to come by, switching vendors may actually be a competitive advantage (if performance/price pans out and you can actually buy the hardware now as opposed to later).
In terms of ML "compilers"/frameworks, afaik there's:
* Google JAX/Tensorflow XLA/MLIR, * OpenAI Triton, * Meta Glow, * Apple PyTorch+Metal fork.
Zen 1 showed that absolute performance is not the end-all metric ( Zen lost on single-core performance vs Intel). A lot of people care for bang-for-buck metric. If AMD can squeak out good-enough drivers for cards with good-enough performance for a TCO[1] significantly lower than NVidia, they break Nvidia's current positive feedback cycle.
1. Initial cost and cooling - I imagine for AI data center usage, opex exceeds capex.
To become a person who writes driver infrastructure for this sort of thing, you need to be a smart person who commits, probably, several of their most productive years to becoming an expert in a particular niche skillset. This only makes sense if you get a job somewhere that has a proven commitment of taking driver work seriously and rewarding it over multiple years.
NVidia is the only company in history that has ever written non-awful drivers, and therefore it's not so implausible to believe that it might be the only company that can ever hire people who write non-awful drivers, and will continue to be the only company that can write non-awful drivers.
The move from PoW to PoS for most crypto networks in combination with bust of ‘22. NVDA slid down in value.
OpenAI debuts ChatGPT in late 2022 and now it’s suddenly bumping in price as the hype and rush for GPUs from companies of all types buys up their stock of GPUs. Demand is far outpacing the supply. Nvda can’t keep up.
Thus, share price is brittle. Competition in the GPU market is dominantly owned by Nvidia. That can change, but so far openai loves using nvidia for some reason.
The compute for a direct answer like that is fractions of a penny, it might be better to create answers on the fly than store an index of every question anyone has asked (well, that's essentially what the weights are after all)
https://www.linkedin.com/pulse/rising-cost-llm-based-search-...
You may wish to look at history to see how things can work out: Cisco had a P/E ratio of 148 in 1999:
* https://www.dividendgrowthinvestor.com/2022/09/cisco-systems...
The share price tanked, but that does not mean that people got bored of the Internet and the need for routers and switches. QCOM had a P/E of 166: did people decide that mobile communications was a fad?
The connection between technological revolutions and financial bubbles dates back to (at least) Canal Mania:
* https://en.wikipedia.org/wiki/Canal_Mania
* https://en.wikipedia.org/wiki/Technological_Revolutions_and_...
It is possible for both AI to be a big thing and for NVDA to drop.
> https://en.wikipedia.org/wiki/Tulip_mania
While widely used as an example, most of the well-known stories about this were actually made up, and it wasn't as bad as it is often made out to be.
Quinn and Turner, when they wrote about bubbles:
* https://www.goodreads.com/book/show/48989633-boom-and-bust
* https://old.reddit.com/r/AskHistorians/comments/i2wfsm/i_am_...
purposefully excluded it because their research found it wasn't actually a thing. (Though for the general public it can be an illustrative parable.)
Competition WILL come. Maybe it's Groq, maybe AMD, maybe Cerebras. Maybe there's a stealth startup out there. Point is, they're going to be challenged soon.
It's almost impossible to manufacture at scale with good yields and leading edge fabs are almost all bought out.
Yes, CUDA, but CUDA is maaaaaybe a few tens of billion USD deep and a few (more) years wide. When the rest of the industry saw compute as a vanity market, that was sufficient. Now, it's a matter of time before margins go to, uhhh, less than 90%.
Does that make shorting a good idea? I wouldn't count on it. The market can always remain irrational longer than you can remain solvent.
Tesla went after them with Dojo and has still ended up splurging on big H100 clusters.
At this point AMD investors should be rebelling, it's pissing money out there but they are not getting wet, and management might have doubled the stock price but that's little consolation if "order of magnitude" is what could have been.
Looking at the chart for $AMD over the past 5 years gives plenty od reasons to be happy, and no reason to rebel. A rational AMD investor should not be Jonesing Nvidia's catching lightning in a bottle via crypto + AI. The Transformers paper was published a few months before AMD released Zen 1 chips - they did not have a lot of money for GPU R&D then.
The timing of the LLM-craze was very fortuitous for Nvidia.
However, given that the nearest competitor AMD has basically given up on building a CUDA alternative, despite the fact that this could grow the company by literal trillions of dollars, I suspect the CUDA moat is much bigger than I give it credit for.
In the past year, they had a revenue of 60B $ and net income of 30B $. Absolutely amazing numbers, I agree. The year before they had a revenue of 30B $ and a net income of 4.5B $ - and it was a rather good year. What happens next of course depend of how you judge the situation - was it a peak hype demand ? Will it stabilize now ? Grow at current extraordinary rates ?
Scenario 1 - margins get back to normal due to hype going down, competition improving etc - in this case the company is worth at best ~200B $ - or 1/10 of what it is now.
Scenario 2 - they maintain current revenue and the exceptional margins - the company would be worth ~1T - or 1/2 of what it is now.
Scenario 3 - they current growth rate (based on past 12 months) continue for ~5 years. This is the case the company is worth ~2T $.
But they are in a business where most money come from a handful of customers, all of which are working on similar chips - and given the sums in play now, the incentives are *very* strong.
My opinion, is that the company is already priced for perfection - basically the current price reflects the perfect scenario. I struggle to see any upside, unless we have AGI in the next 5 years and it decides it can only run on Nvidia chips.
All of this is akin to Tesla in the past years. They grew from a small startup to a medium car maker - the % growth rate was huge of course - an amazing achievement in itself. But people projected that the % growth rate would continue - and the stock was priced accordingly. Reality is catching up on Tesla, even if some projections are still absolutely crazy.
If you're convinced the stock is that overvalued, go short some or, if you like to live dangerously, buy some long-term put options (don't be an idiot and buy short-term options.)
I have no idea if NVDA is like Cisco Systems in 2000, or if it's something unique. What I am aware of is that there's around 5-7 trillion that were moved from stocks to t-bills since the Fed raised rates in March 2022. If and when they drop their rates back to the historical ~2.5%, it's reasonable to predict these funds will go back into stocks, which will presumably drive up prices.
Buying long-term put options on Nvidia now is extremely expensive - the stock was so volatile that the price you pay for those options almost annihilate any gains you could expect, even if the stock losses 50% in 12 months.
You got me curious about those 5-7 trillions. Where these numbers come from ?
Frankly, I don't understand why we made it possible for individuals to gamble by selling options. As Charlie Munger used to say, Wall Street will sell shit as long as shit can be sold.
Selling cash secured puts or selling covered calls would be less risky than just holding stock.
For us, older folks, we've seen this 'new normal' several times already - it will end up as usual. There are no free lunches and as it appears to me that have not entered any permanently high plateau.
It's even quite funny that ~100 years ago, we've had the previous big pandemic, and the biggest stock market crash. Epidemic of this century is done, now waiting for the second part !
To clarify, I’m not saying NVDA won’t crash from here or bear markets no longer exist. I’m simply saying that historic PE valuations are a poor metric for assessing the potential of a stock in todays market conditions.
Yes, it's probably the first time that retail is allowed to trade options. But it's not the first time that retail is all in in stocks. I've tried to find a funny number to back it up - just check the Wiki on 1929 crash - https://en.wikipedia.org/wiki/Wall_Street_Crash_of_1929 - there was more money lent to 'small investors' so they can buy on margin ... than the entire amount of currency circulating at the time.
On average, all those retail guys will loose money - that's the sad truth. In the long term, the stocks simply follow the earnings - all other movements around this trends are pretty much a zero sum game - and most skilled operators are not loosing money in that game.
Price to earning ratio is just the number of years the company 'pays for itself' if you buy it. PER at 40s for big chunks of main indexes mean that either there will be tremendous progress in the economy that will boost the earnings or people are hoping to resell to a bigger fool.
Note: I work in finance, and I very much see the retail involvement in stocks. Hedge Funds and banks, are making a ton of money out of them, that's for sure.
You've also had three or four bona fide bubbles in that span, starting around 2017. First was Bitcoin along with the stock market as a whole (with Nvidia being one of the leading stocks of that bull market advance).
Then you had Tesla go parabolic and lots of people become rich. Then you had the whole post-COVID speculative mania.
The result of this has been extreme credulity by the average person. Today's keynote is the perfect summation of this phenomenon. I saw multiple people who almost certainly couldn't explain in any level of detail how Nvidia GPUs are used for training and inference, but rather rely on the secondhand talking points like CUDA that they've learned by watching Jim Cramer, watching this keynote with excitement and anticipating how much it would pump their shares or call options.
Contrast this with Steve Jobs keynotes from 15 years ago when Apple's best days were well ahead of them. Most keynotes were questioned, in some cases even mocked. When Tesla stock broke out, many people couldn't make sense of it. Ditto for cryptocurrencies. But now, taking their cues from those cycles, the average person wants to ride the next bubble to riches and is trying to catch the wave and so now believes every story attached to a rising asset price.
CEO's aren't blind to this and are using every opportunity to create favorable storylines. The leadup to a keynote like this carries with it an enormous amount of pressure to deliver. Hence a company like Nvidia leaning into generative artwork and straight up made up storylines like robot development.
At the end of the day, I'm afraid that there likely isn't all that much substance and the evidence is beginning to pile up that the megacap tech stocks have run out of ideas which is why they are laying off people en masse and appealing to the AI hype cycle to carry their stocks higher.
Consider that Nvidia has gone up 8x- 800%!- in just over a year. The cycles are moving faster and faster. I remember just a year ago when lots of people said Nvidia at $250 was insane. Now here we are with the stock at more than three times that level and most people are calling it cheap. The stock market seems to have in certain areas like semis, completely disconnected from the fundamentals and taken flight. Yes, Nvidia earnings have grown. But understand that this is all part of a positive feedback loop where tech CEO's are pressured by their competitors and shareholders to show that they are investing in AI. Thus they all talk about it on their earnings calls and spend massively. All of their stocks rise in unison as you have a market that increasingly looks like its chasing momentum stock trends up. Nvidia's moves of late have almost nothing to do with any fundamental developments in the company. It has been routinely trading upwards of $45 billion a day. The Friday before last that number was over $100 billion. These are absolutely insane figures. Compare that to Microsoft, the largest company by market cap in the world, which trades on average around $8 billion per day.
I think this is generally how bull markets end and I think we may be actually forming the top of the great bull market for the megacaps that began around 2010 but really hit its stride starting in 2017.
However they went up 8x because (neglecting crypto) they overnight transitioned from providing accessories to PC gamers and high end engineering workstations (both increasingly niche markets with tapering growth or decline) to being for the moment the only substrate of an entirely new consumer product segment that has seen the most rapid adoption of any new technology in the history of the world.
This could be the way things work now: the time constants shrink as the pipeline efficiency increases.
This doesn't make much financial sense. Since every T-Bill is held by someone the amount of T-Bill outstanding is completely determined by government issuance. And since the amount of T-Bill outstanding will only grow over time, no "money" will ever flow out of it. Now if you narrow you inclusion of T-Bill holders to a specific group of people the amount this group holds could certainly go up and down over time. But then I wonder how you know this group sold stocks to buy T-Bills, and why this is more significant than the action of their counterparties: every share they sold was bought by someone else after all.
They are waiting for earnings projections for such a pop since right now it is extremely overbought and struggling to move past >$1,000 per share.
For now, Microsoft and OpenAI will use these chips, but in the long term they are just looking at this and plotting to build their own chips and reducing their dependence on Nvidia and will be ready to switch once their contracts have run out.
I think there may be a typo though, I assume this also includes liquid-cooled vs air-cooled.
[1] https://nvdam.widen.net/s/xqt56dflgh/nvidia-blackwell-archit...
Seems revenue from inference is growing at a significant clip.
though it seems most of the progress has been on memory throughput and power use which is still very impressive.
I wonder how this will trickle down to the consumer segment.
It doesn't matter for anyone who's not microsoft, aws or openai or similar.
Again, I do think the throughput and energy efficiency gains are impressive, but the raw performance gain is lower than I'd have expected for the massive leap in node size etc
It's basically TensorRT-LLM + Triton Inference Server + pre-build of models to TensorRT-LLM engines + packaging + what appears to be an OpenAI compatible API router in front of all of it + other "enterprise" management and deployment tools.
This software stack is extremely performant and very flexible, I've noted here before it's what many large-scale hosted inference providers are already using (Amazon, Cloudflare, Mistral, etc).
From the article:
'Nvidia will work with AI companies like Microsoft or Hugging Face to ensure their AI models are tuned to run on all compatible Nvidia chips. Then, using a NIM, developers can efficiently run the model on their own servers or cloud-based Nvidia servers without a lengthy configuration process.
“In my code, where I was calling into OpenAI, I will replace one line of code to point it to this NIM that I got from Nvidia instead,” Das said.'
The dead giveaway is "I changed one line of code in my OpenAI code" which means "I pointed the OpenAI API base URL to an OpenAI compatible API proxy that likely interfaces with Triton on the backend via its gRPC protocol".
I have a lot of experience with TensorRT-LLM + Triton and have been working on a highly performant rust-based open source project for the OpenAI compatible API and routing portion[0].
On this hardware (FP4) with this software package 30x compared to other solutions (who knows what - base transformers?) on Hopper seems possible. TensorRT-LLM and Triton can already do FP8 on Hopper and as noted the performance is impressive.
[1]https://twitter.com/swyx/status/1760065636410274162?t=rpbcr8...
It’s two fused chips. So 1.25x per chip. 25% uplift. Not 2.5x uplift. The 2.5x is for the whole package.
Then the B200 package is 2 of these plus a CPU. So a total of 4 GPU dies in each unit.
That's GB200.
Jensen's comment about being first was such a dig to Emerald Rapids.
I believe this is just the first for a GPU only product.
Dell proves that selling complete units is very profitable.
Apple shows that owning the entire stack is immensely profitable.
Nvidia already has significant hardware and software investment. They very well could fully integrate and grab larger slices of the pie.
In fact, Nvidia already has complete appliance like fully integrated machines. But enterprises like to install their own OS and run their own software stack. These appliances have not caught on, at least not yet.
Apple shows no such thing. Apple, sells pretty, reliable and safe. A car is a car, but apple is a sports car, or a saloon. Vertical integration is the way they chose to deliver that, and pretty and reliable are all normal people care about.
Nvidia is gonna have to think long and hard about the "whole stack". 20 years ago they might have been able to pull a next, but right now anything that isnt LINUX is a rounding error, and they dont want to turn into SUN (no one at nividia is smart enough to make them the next sun).
Nvidia architecture + Nvidia os is not something that I see them pulling off for the datacenter.
"Ultimate AI Workstation". Pricing starts at US$83,549:
https://shop.lambdalabs.com/gpu-workstations/vector/customiz...
Adding every option only adds $2,100 to the price too (totalling $85,649). They should probably just include everything as standard. ;)
Making NNs generalize the way humans do it is still a hard problem.
Then he went into a comedic monologue about GANS. But hey, at least that meant that the CEO was reading the actual conference proceedings...
Don’t think anyone cares as long as the company keeps getting the bets right
Nvidia is not becoming a cloud provider (beyond a small eval environment perhaps).
https://developer.nvidia.com/blog/nvidia-nim-offers-optimize...
so there is a push towards a platform company, but probably not well explained in this article.
Platform company, as in they're allowing developers on their AI platform and opening an app store?
The article doesn’t discuss becoming a platform co but instead discussed ways their existing platform subscription model is evolving to add backwards compatibility testing.
So far the short sellers, learned that bitter lesson.