This particular hype train has about 2-3 years more of gas until Microsoft/OpenAI and Meta stop buying absurd amounts of other peoples' gear and design their own.
This particular hype train has about 2-3 years more of gas until Microsoft/OpenAI and Meta stop buying absurd amounts of other peoples' gear and design their own.
Microsoft's and Meta's weaknesses are that they don't directly have a specific need for all of this hardware. Much of it is being consumed by R&D without any products or profits involved directly. At some point, the pressures to monetize will come calling and turn off the blank check spigots.
Right, and they are standardizing on CUDA. Something no other company has any chance of matching for the forseeable future.
How many people are writing to CUDA directly, and how many folks are using something like PyTorch front-end where the back-end can potentially be changed?
* https://pytorch.org/blog/pytorch-for-amd-rocm-platform-now-a...
* https://rocm.docs.amd.com/projects/install-on-linux/en/devel...
* https://dev-discuss.pytorch.org/t/opencl-backend-important-u...
And even with middleware, turns out CUDA has the best debugging tooling, followed by DirectCompute/DirectML and Metal.
now, for how many of those things does AMD have a pluggable backend that actually works? none of them, unfortunately. Octane implemented on SPIR-V - on NVIDIA and Apple, because those were the runtimes that worked. AMD and Intel couldn't even compile it successfully (intel probably would today, I'd guess). Blender tried to support AMD for a long time, and the runtime was so buggy and defective they eventually pulled support from OpenCL entirely, since the OpenCL only ever ran properly on NVIDIA and Apple, and both of those have alternative interfaces that are cleaner and faster. The only point was supporting AMD, and AMD's OpenCL implementation is so buggy that it ended up being a de-facto AMD-specific codepath anyway.
That's the problem, AMD's software stacks are fractally broken. Not just HIP and their oneAPI adapter, but the ROCm it runs on. Not just ROCm, but the OpenCL it runs on. Not just openCL, but the kernel scheduler it runs on. Everything is broken at every level even on supported hardware.
Saying "it's just software" is a handwave. There are literally entire stacks here that are not fit-for-purpose and need to be substantially rebuilt before you can start building the layers above them. There isn't a stable enough basis to just implement an adapter for these pluggable frameworks, because the runtime and the kernel code are also broken.
there's a fairly incredible amount of commercial value in a system that can provide a workable "fuzzy match" or a system that can produce an approximately-optimal outcome without an intractable optimization step, for example. what is the commercial value of making the entirety of US logistics even 5-10% more efficient, or work-scheduling optimizations in a server farm?
like people just keep blindly repeating that there's no profitability there regardless of all the places it's quietly being deployed to great effect.
IIUC, OpenAI leased the model to Microsoft to power LLM search on Bing. OpenAI doesn't run the inference on their machines. Microsoft also added LLM features into Office (and Windows, too, right?)
What do you think is powering all that inference?
How many times in history has this ever happened? That a product was so good no competitors could match it, forever?
A better example might be Intel, but Intel at least had a unique instruction set. And of course Intel is more of a cautionary tale than a story of eternal market dominance.
Analogy: Apple's hardware + software combination as competitive advantage
Are they just the beneficiaries of inside (R&D) or outside (AI) market forces?
History has shown that the latter is not sustainable (dot-com, nanotech, covid, crypto, AI)
TLDR: Shovels eventually become commodities
(although I agree with the point that yes, NVIDIA was thinking about systems engineering much sooner and more comprehensively than AMD was. And not just multi-node scaling, but also things like the programming model (CUDA vs OpenCL/AMD APP/HSA Framework), and they also seem to have correctly identified that ethernet allows higher bandwidth-density than pcie as a transport for this interconnect, giving them an implementation advantage. Similarly, despite all the praise for AMD's packaging, it's often come with downsides and caveats, and NVIDIA usually isn't far behind with one that doesn't. Eg Fury X/Vega/Radeon VII vs GP100/GV100/GA100, or MI300X vs B200. And NVIDIA has explicitly been much earlier-to-market on the systems-scaling stuff like NVSwitch or pre-specced high-performance systems modules like DGX, which AMD still doesn't really have an answer for despite NVSwitch being over a decade old at this point etc.)
Tesla built a whole interconnect system for Dojo too, iirc, that was an explicit focus in their design. Not surprising for Jim Keller. But Jensen Huang nailed that requirement too.
AMD is definitely behind and slower than I would think they should be, but they are making progress. I suspect they are working pretty seriously on it, just doing so quietly. At least, I hope that's what's going on because if not, it's a tragedy.
I'm pretty sure it's been more than 5 years. People have been griping since before ROCm was released which was around 7 years ago.
TSMC revenue is smaller than Nvidia revenue.
TSMC profit margin is great, but not as great as Nvidia.
N.B.: there are other 3rd-party competitors like Cerebras [2] who offer all-in-one solutions for their giant wafers along with libraries and data centers but I’m not sure behemoths would migrate to these offerings either
[1] https://ai.meta.com/blog/next-generation-meta-training-infer...
I doubt they'll continue to buy a significant amount of GPUs if they stop seeing significant enough performance/quality gains.
I'm also not clear how much of the cost of running models is really because the models are too large to fit on a single device? If we start making devices with very large (terabytes) of memory, how much compute do we then need?
tldr; Would it not be cheaper to build more specialized hardware that has massive amounts of memory (enough to hold the entire model on a single device)? I'm wondering if this sort of hardware (large banks of memory + relatively simple compute) is not a market that many of NVIDIA's competitors can easily enter.
That's a gamble. It's the whole idea behind a pre-trained model (GPT). If you make a really good model then you can use it a lot of times. That seems to be playing out as intended at the moment.
> very very very big matrix operation
To do efficiently, it's many slightly smaller very very big matrix operations.
> does it require advanced GPU capabilities
No, but GPUs don't have a lot of overhead for those very very big matrix operations.
> high costs because models are too large to fit on a single device
Engineer costs are high, and communication costs are also high. Those are mostly just barriers to entry though. Compute still dominates.
> competitors with more RAM
It's a little more complicated than that because extra RAM has speed-of-light delays and other such problems. If you don't expose a CUDA interface then you have a long hill to climb to get even a couple customers. A couple extra caches or other more esoteric memory APIs won't be easy for the industry to adopt (maybe a clever electrical engineer can fix that; I don't have hard proofs of those bounds).
no, training takes hundreds/thousands of high-end HPC gpus for days/weeks/months/years. Inference runs in seconds.
The argument around "inference is more important/will be a larger market" revolves around the total volume of inference performed being larger than the amount to train. A million consumers perform inference, but you only train it once.
The problem is the training really never stops, it's not like you train a model and you're done. Even if there were no new innovations it would be worth continuing to train models for years even just fitting the models we've got to the data better. LORAs and QLORAs to train. Etc. As you adopt it more widely, the amount of things we need to train will continue to increase anyway. There is no particular reason to assume the growth of training needs will ever stop, or even decrease below current growth rates (TOPS, not $). It will certainly not continue at 300% margins forever, of course. But I just don't think the assertion that training needs will decrease is supportable - this seems like a Jevon's Paradox moment, the increased utility and accessibility of this will drive an enormous increase in usage, both training and inference.
It's also a misleading statement around revenue. Yes, there will be a lot of inference that's done... and it's trivial to implement and will be in almost every single device etc. Training is really the only thing it's feasible to build an ecosystem/moat around, because inference is going to be done by everyone, and will have almost zero margin for either standalone devices or the IP. Training is the interesting and profitable part of the market, and new discoveries go from a "research -> training/scale-out -> inference" pipeline that tends to favor NVIDIA because they have the ecosystem for the research etc. It has been the flexibility and programmability of the GPGPU model that has won so far - you can certainly get a lot of TOPS on Trainium or whatever, but if they can't train the techniques currently used then it doesn't matter.
They have spent an unfathomable sum of money (tens of billions) fostering that ecosystem, sponsoring academics and researchers and conferences and paying for people to write libraries and tuning their software etc. There are millions of man-hours of work that need to be done even if you fully understand exactly the minimal combination of pieces to get there, and that doesn't even begin to replicate the mindshare factors etc. Intel increasingly has a perfectly viable software ecosystem, but that doesn't matter if nobody is using it. NVIDIA and Apple are really the only 2 ecosystems anyone uses, in the sense of being a place where work takes place on the platform for the sake of being used by people on that platform. Everything else is just a port-of-convenience for access to hardware and nobody cares a whit about building a long-term community around SyCL vs ROCm as a thing in itself. That's table stakes to build an ecosystem, not the ecosystem itself. Having a complete, viable software ecosystem was the ask 10 years ago, to be competitive in 2024 you need the people using it because we're no longer in the blue-sky/academic-research stage here.
It's like thinking you can be a reddit competitor just because you implemented lemmy - the thing that makes reddit reddit is not the source code, that's table stakes, not a platform. Building a mapping layer over the top might help, but it also removes any incentive for the meta-layer to care about you as anything more than a "utility". Literally the players that are going to succeed are the ones like NVIDIA, Apple, and Sony where there is a first-party culture and an ecosystem that is intellectually self-sustaining as a target in itself (and not a utility/build target for someone else). Maybe microsoft to a lesser extent (they are in a good place given their control of the dominant client OS).
The "iphone bubble" never popped - yearly revenue increased more or less continuously for a decade, even during the Great Recession it went up. Even though the smartphone market did see many other viable players emerge, Apple retained a plurality control of the market and remains the single largest and most influential player in design as well as direct market influence. Saying something is a "bubble" is begging the question, maybe NVIDIA just popularized a whole new device/field (GPGPU) in the same way Apple did and the market grows underneath them. People are blindly making the assertions that the revenue MUST be unsustainable because "well I don't see any value!" and then building this whole little chain of logic on some things that are ultimately propositions/assertions and not fait-accompli.
The moment it happens, Nvidia will stop being that profitable from hardware sales there. For now they should be able to break records YoY. And indeed it feels like 2-3 years it will last.
Intel and Microsoft had the problem of competing with their previously sold products. Most of their users only needed a basic PC for email, Word, and Excel. Nvidia doesn't have this problem because previous hardware generations quickly become obsolete as the demand for compute keeps growing and it's easier to manage newer clusters.
Fixed that for you.