An Interview with Nvidia CEO Jensen Huang About AI’s iPhone Moment
stratechery.com
stratechery.com
- Apple have got their own chips now for mobile AI applications for iPhones.
- Google have got their own chips for Android and servers.
- AMD are gaining on Intel in the server space, and have their own GPUs, being able to sell CPUs and GPUs that complement each other may be a good strategy, plus AMD have plenty of their own experience with OpenCL.
Google has TPUs but have these even made a tiny dent in Nvidia's position?
I assume anything Apple is cooking is using Nvidia in the server room already
Intel seems completely absent from this market
AMD seems content to limit its ambitions to punching Intel
its Nvidia's game to lose at this point...I wonder when they start moving in the other direction and realize they have the power to introduce their own client platform (I secretly wish they would try to mainstream a linux laptop running on Nvidia ARM but obviously this is just a fantasy)
if anything, I think Huang may not be ambitious enough!
Nvidia is making the right call, of course.
I fully expect future rendering techniques to lean heavily on AI for the final scene. NeRF, diffusion models, et cetera are the thin end of the wedge.
Yeah. It's good hardware. You can get cheaper cards (even cost-competitive options) on PC but Nvidia won't sell them to you. Especially not now that they're got 10 billion dollars on their TSMC tab.
Nvidia will not give up gaming. When every gamer has a Nvidia card, every potential AI developer to spring up from those gamers, will use Nvidia by default. It also helps gaming GPUs are still lucrative.
That's a great counter point.
> Nvidia will not give up gaming. When every gamer has a Nvidia card, every potential AI developer to spring up from those gamers, will use Nvidia by default. It also helps gaming GPUs are still lucrative.
Another. But Nvidia will have a lot of balancing to do and some very thirsty competitors. Though if competition arises, that too is good for gamers.
I wouldn't be so quick at assuming this. Apple already ship ML-capable chips in consumer products, and they've designed and built revolutionary CPUs in modern time. I'm of course not sure about it, but I have a feeling they are gonna introduce something that kicks up the notch on the ML side sooner or later, the foundation for doing something like that is already in place.
Mac Minis don't count
Sure, they are super rich and could just buy their way into the space...but so far they are really far behind in all things AI with Siri being a punchline at this point
if anything, Apple proves that money alone isn't enough
The iPhone was their first phone, and it really kicked in the smartphone race into high gear. Same for the Apple Silicon processor. And those are just two relatively recent examples.
So both are not very good examples, because they build up experience over long periods.
They are examples of something they could similarly do for the Apple Neural Engine but in a bigger scale in the future. They have experience deploying it in a smaller scale/different versions, they would just have to apply it in bigger scale in order to be able to compete with NVIDIA.
Also, while Apple did create their first chip (at least of their current families) in 2007, they did acquire 150 or so engineers when they bought PA Semi in 2008. So, that gave them a leg up compared to building a chip team completely from scratch.
I doubt Apple cares about spending a few hundred million dollars on A100s as long as they make sure the resulting models run on billions of apple silicone chips.
Has Nvidia not done that too? They shipped ML-capable consumer hardware before Apple, and have revolutionary SOCs of their own. On top of that, they have a working relationship with the server/datacenter market (something Apple burned) and a team of researchers that basically wrote the rulebook on modern text and image generation. Then you factor in CUDA's ubiquity - it runs in cars, your desktop, your server, your Nintendo Switch - Nvidia is terrifying right now.
If the rest of your argument is a feeling that Apple will turn the tables, I'm not sure I can entertain that polemic. Apple straight-up doesn't compete in the same market segment as Nvidia anymore. They cannot release something that seriously threatens their bottom line.
If they manage to move a significant part of ML compute from datacenter to on-device, and if others follow, that might hurt Nvidia's bottom line. Big if at this point, but not unthinkable.
Then there's the issue of model size. You can fit some pruned models on an iPhone, but it's safe to say the majority of research and development is going to happen on easily provisionable hardware running something standard like Linux or FreeBSD.
And all this is ignoring the little things, too; training will still happen in-server, and the CDN required to distribute these models to a hundred million iPhone users is not priced attractively. I stand by what I said - Apple forced themselves into a different lane, and now Nvidia is taking advantage of it. Unless they intend to reverse their stance on FOSS and patch up their burned bridges with the community, Apple will get booted out of the datacenter like they did with Xserve.
I'm not against a decent Nvidia competitor (AMD is amazing) but the game is on lock right now. It would take a fundamental shift in computing to unseat them, and AI is the shift Nvidia's prepared for.
For training, sure. For inference, Apple has been in a solid competitive position since M1. LLaMa, Stable Diffusion, etc, can all run on consumer devices that my tech-illiterate parents might own.
This seems unknowable without Google's internal data. The salient question is: "how many Nvidia GPUs would Google have bought if they didn't have TPUs?"
The answer is probably "a lot", but realistically we don't know how many TPUs are deployed internally and how many Nvidia GPUs it displaced.
https://www.nextplatform.com/2022/08/23/inside-teslas-innova...
Much like Google, I think Tesla realized this is a capability they need, and at the scales they expect to operate, it's cheaper than buying a whole bunch of NVIDIA product.
What's the deal with that anyway? A lot of people want a real alternative to Nvidia, and AMD just... Doesn't care?
I guess we'll have to wait for intel to release something like CUDA and then AMD will finally do something about the GPGPU demand.
When AMD bought ATI they viewed the GPU as a potential differentiator on CPUs. They've invested a lot of effort into CPU-GPU fusion with their APU products. That has the potential to start paying off in a big way sometime - especially if they figure our how to fuse high end GPU and CPU and just offer a GPGPU chip to everyone. I can see why AMD might put their bets here.
But the trade off was that Nvidia put a lot of effort in doing linear algebra quickly and easily on their GPUs and AMD doesn't have a response to that. Especially since they probably strategised on BLAS on an APU. But it turns out there were a lot of benefits to fast BLAS and Nvidia is making all the money from that.
In short, Nvidia solved a simpler problem that turned out to be really valuable, it would take AMD a long time to organise to do the same thing and it may be a misfit in their strategy. Hence ROCm sucks and I'm not part of the machine learning revolution. :(
https://www.pcgamesn.com/amd-sony-ps5-navi-affected-vega
https://www.pcgamesn.com/amd/rdna-2-sony-ps5-gpu-pc
As such, if the console market doesn't want it, it doesn't get built. AMD is not willing to put its own money into graphics research.
AMD does not really have the marketshare to get the PC market to adopt AMD-backed features that use accelerators that aren't present in the consoles. If AMD takes 20% of the market in a given year, and the PC market turns over every 6 years, this hardware support would be present in 0-3% of the PC market and 0% of the console market. So even if RDNA3 had a magic "DLSS-level" improvement that relied on some unique new accelerator they'd added in RDNA3, it'd be an uphill fight to get it adopted. Nor is AMD going to spend the money to just implement a bunch of software features anyway - they only even invested in FSR2 after it became a competitive disadvantage for them not to have something.
They won't even go the 16-series vs 20-series route of having consoles be a basic architecture (with size-reduced implementations of features) and then a full-size/higher-performance implementations on PC dGPUs with more full-fledged accelerators bolted on/etc. For example they could have done this with the ML accelerators on RDNA3 - they have a slower (microcoded?) ML instruction in the basic RDNA3, and they could have thrown a more full-fledged implementation into dGPU implementations where there's more space to spare.
But it's just not worth spending on any of that for them - it's a lot of R&D for a fairly narrow slice of the market that would be impacted.
https://www.anandtech.com/show/13973/nvidia-gtx-1660-ti-revi...
So yeah I mean she's just not that into you. Consoles set the direction of their graphics R&D. They'll tap a few other lucrative markets like HPC but they're not going to make big spends that don't have obvious ROI involved, and AMD doesn't really have the PC-gaming marketshare to care about dGPUs as an independent market worthy of R&D.
People ask "why does Intel need anything except iGPUs" and for AMD the question is "why do they need anything except consoles". The rest is interesting in a "someday" sense and potentially strategically important, but day-to-day it's pretty obvious which verticals are bringing in the bacon.
And for NVIDIA that's both dGPUs and datacenter - they still make a lot of money from consumer gaming, and it gives a foothold for development to progress from curiosity to research project to business deployment. AI accelerators and CUDA being on consumer hardware has been a huge boon to R&D (contrast ROCm/HIP being essentially unusable outside enterprise hardware) and the commercial market has found uses for RT cores as well. Because NVIDIA had the realization, a lot of years ago, that they are in fact a software company, that writes the software that sells the hardware.
People mocked Jensen for that for a lot of years, but he was completely right and that's why he's succeeded while AMD has spun their wheels on GPGPU for 15 years now.
And the problem for AMD is, consoles won't pay for a 5% more expensive chip based on blue-sky prospects of something maybe being useful in 3+ years. Or at least not unless it gets an internal backer, like DirectStorage/RDMA obviously has been adopted despite an extremely slow burn on actual usage.
Optical Flow Accelerator is probably the most recent iteration of this - GCN actually had this capability as "Fluid Motion" accelerator but consoles wanted it taken back out, because it was wasted space. Now it's the underpinning of DLSS3 and likely future work in DLSS4 - the principles of "variable temporal+spatial rate shading" AMD outlines in their recent GTC presentation seem like an obvious "DLSS2 for DLSS3". I have also spoken about this idea before and I think that is where NVIDIA is going with DLSS4, but AMD has to do it without the hardware optical flow engine (except on older GCN cards ironically).
https://gpuopen.com/gdc-presentations/2023/GDC-2023-Temporal...
I don't think Apple's server side is big or interesting. Far more interesting is the client side, because it's 1bn devices, and they all run custom Apple silicon for this. Similarly Google has Tensor chips in end user devices.
Nvidia doesn't have a story for edge devices like that, and that could be the biggest issue here for them.
Tangent but I wish they would!
Apple's e-cores would be great for servers, and they are very area/transistor efficient (even considering the node). 0.69mm2 for something with (broad strokes) Gracemont-ish performance/skylake-ish performance (but no SMT) is really good even considering the 5nm node.
I think the real-world density shrinks on 5nm ended up being around 60%... so 0.69mm2 for Blizzard is like 1.1mm2 equivalent on 7nm and Avalanche is 4.1mm2, versus Zen3 at 3.1mm2 and Zen2 at 2.72mm2.
10ESF density is supposed to be similar to TSMC 5nm, dunno how true that really is in practice on actual products. But on paper that means you have Gracemont at 1.7mm vs Blizzard at 0.69mm2 and Golden Cove at 5.55mm2 vs Avalanche at 2.55mm2.
Or comparing to AMD using the 1.6x conversion factor, that gives you a 7nm-area-equivalent (assuming 5nm density on 10ESF) of 2.72mm2 for Gracemont (vs Zen2 at 2.72mm2) and 8.88mm2 versus 3.1mm2 for Zen2. And that's why they're doing e-cores, and AMD is just squeezing the last little bit of space out of their existing uarch, lol.
https://www.reddit.com/r/hardware/comments/qlcptr/m1_pro_10c...
The M1 Pro/Max dies are mostly consumed by a gigantic iGPU (in a way it's similar to the latter days of Intel quadcore era) but the cores themselves are actually quite svelte - it's actually not a case of Apple "just throwing more transistors at it", sure they are doing that in the GPU but the CPU cores themselves are very area-efficient (again, even considering the node).
https://en.wikichip.org/wiki/File:kaby_lake_(dual_core)_(ann...
https://en.wikichip.org/wiki/File:kaby_lake_r_die_shot_(anno...
A Sierra Forest-style product with multiple chiplets full of nothing but e-cores would be a fantastic thing. I completely agree that Apple doesn't have any notable presence in server, but, you could make some real good products with the pieces Apple has already demonstrated.
I don't have an exact source, but I recall the Asahi folks saying that based on their reverse engineering, Ultra/2-chiplets isn't the limit, the architecture is laid out to go higher on chiplets (I want to say 4 or 8) and they just aren't exploiting it right now.
...but, if core size is your jam (for whatever reason), keep an eye on Nvidia's Grace CPU. It's their stab at a datacenter-scale ARM SOC, and it should be releasing before EOY. Then there's the Ampere offerings that already have acceleration for PyTorch, ONNX and Tensorflow, along with Graviton for general-purpose efficiency... there's a lot of low-profile ARM cores in the datacenter today.
A good start for Apple would be updating the rackmount Mac Pro with an 80 core Double Ultra chip, but even that feels fairly pedestrian next to the 144-core-complex Grace is teasing. I'm sure it sounds silly to the readers of this website, but I genuinely don't think Apple is up to the task of competing in the datacenter. Obviously so on the software side, but arguably not even on the hardware front either.
I don't see what Apple stands to gain from getting into the enterprise market, other than simple diversification of their portfolio.
They'd stand to win a market they have no foothold in and profit.
On the other hand, Apple's non-existence in the enterprise space isn't due to a lack of trying. They've been there and done that already.
Imagine if AMD launch a new bus implementation from CPU to GPU. That's not something Nvidia can do by themselves. Maybe Nvidia buys Intel and does it though!
https://www.tomshardware.com/news/amd-infinity-fabric-3
AMD already does that, and much like G-Sync there is also an open standard (CXL) that everyone else is converging around.
It's indeed ages beyond any of their competitors. However, most ML/DS people interact with CUDA via a higher-level framework. In recent years this community has consolidated around a few (and even only one platform, PyTorch) framework. For some reason AMD had not invested in platform backends, but there is no network effect or a vendor lock-in to hinder a shift from CUDA to ROCm if it is supported equally well.
AI hardware seems like it can be much simpler than GPGPUs, given the successful implementations by many companies including small startups.
AI hardware software seems like it is extremely difficult. Making a simple programmer and ops interface over a massively parallel, distributed, memory bandwidth constrained system that needs to be able to compile and run high performance custom code out of customers' shifting piles of random python packages.
AMD has continuously struggled at (2) and hasn't seemed to recommit to doing it properly. AMD certainly has silicon design expertise, but given (1) I don't think that is enough.
Xilinix is in interesting alternative path for products or improving AMDs software/devex. I'm not sure what to expect from that, yet.
What chips for Android?
At the same time, it doesn't seem like a great moat - I think AMD should be to able compete pretty well soon. I think TSMC/Samsung/ASML will capture quite a lot of the profit in the boom to come.
Nvidia will have a first-mover advantage, but AMD will be strongly motivated to win those customers, and they should be able to, eventually.
Anyway, that's my gut instinct based on how these things tend to play out. Now let's hear from e.g. actual experts in the field? :)
The better CUDA is, the more things get built on top of it.
The more things get built on top of CUDA, the better the CUDA ecosystem of tooling, frameworks, etc. gets.
The better the CUDA ecosystem is, the better CUDA is. GOTO 1.
But you're right -- lots and lots of folks will be motivated to compete.
Is that also like OpenCL vs CUDA in reverse?
Nvidia is a lonely proprietary ship in a hostile sea of open source packages. It might take 20 years if Intel and AMD fail to do anything useful, but sooner or later they'll either open source CUDA defensively or get ground down by an open library. There are a finite list of features to implement before "Runs on Nvidia" vs "Runs on any graphics card" becomes the only important box left on the bureaucrat's checklist.
The math here is literally 1st & 2nd year university subjects. That is a good short term moat, but not a defensible one long term.
AMD has been trying to catch up in the machine learning space for many years now, and it hasn't happened. For things to finally improve, I guess something needs to change internally at AMD. Even Apple Silicon, in the short time it has existed, has gained better support for ML (e.g. compare running PyTorch on Apple Silicon vs running PyTorch on AMD).
You're not wrong, it's just a fun sign of how far the cryptocurrency space has come.
> the market just wants the software and if the software works the same, the hardware is replaceable
And likely to be true given how much competition is heating up in the AI hardware space. Granted, many of these competitors and startups especially have existed for years and haven’t made much of a dent. Even Google’s TPU doesn’t seem that much better than Nvidia’s stuff based on their limited MLPerf score releases. Maybe this “iPhone moment” for AI will change that and force competitors to finally put some real effort in it.
As for Nvidia, looks like they are trying to adapt by selling their own enterprise software solutions such as Omniverse and their AI model customization stuff. Will be interesting to see if they can transform into more of a software solutions provider going forward.
Also segmenting gaming GPUs from AI GPUs is just not the best idea: devs don't want to buy a gaming GPU and an AI GPU separately when they can just buy a stronger NVIDIA GPU that does both.
"Inference will be the way software is operated in the future. Inference is simply a piece of software that was written by a computer instead of a piece of software that was written by a human and every computer will just run inference someday."The two Acquired episodes about Nvidia are fantastic and worth listening to:
Nvidia: The GPU Company (1993-2006): https://www.acquired.fm/episodes/nvidia-the-gpu-company-1993...
Nvidia: The Machine Learning Company (2006-2022): https://www.acquired.fm/episodes/nvidia-the-machine-learning...
More simply, when every leap in compute is a massively parallel architecture, every problem seems like it needs to be solved by a massively parallel system. But I'm guessing before this cycle is over we start to see the limitations of that.