Google Launches AI Supercomputer Powered by Nvidia H100 GPUs
tomshardware.com
tomshardware.com
They've been trying to push their GPUs as CPU alternatives everywhere especially in the datacenters where their presence grew since the acquisition of Mellanox. They also tried to acquire ARM, to squeeze both Intel and AMD out of the CPU market completely.
I hate what they've done to the PC gamers, but as a company trying to grow in more markets and make even more money, they've executed insanely well strategically, leaps ahead of AMD.
Usually when I see this, it symptomatic of major organizational dysfunction of some type. One time, I was in a firm where 100% of the energy was spent on quarterly objectives for some executive bonus pay structure. No one cared if the organization lived or died, since there were always other jobs.
Nothing AMD is doing in the GPU space is aligned with long-term survival or competitiveness.
Nvidia are effortlessly crushing AMD right now and as far as I can tell it is because they implemented a bunch of BLAS functions on the GPU (it is weirdly difficult to get a good tutorial on how to do matrix multiplication on an AMD GPU; every so often I look for one and have I think literally never found an example). But strategically, AMDs approach to GPU-CPU memory fusion is probably going to be the technically stronger approach. Assuming it works.
In hindsight they should have focused on libraries to let people use their GPU, but big picture they clearly understand how important it is to embrace general purpose compute and are treating it as a high priority.
I mean if anything, Nvidia is already there and crushing it too. CUDA has a unified memory model on Linux today and has for years, so if you have a proper pointer created by cudaMallocManaged, it can be used transparently in both GPU and CPU code without cudaMemcpy. And on the Grace Hopper chip, the open-source driver supports heterogeneous memory management, giving both the CPU and GPU unified, coherent memory across the CPU and GPU even though they have completely separate and isolated memory chips; 512GB LPDDR5X versus 96GB HBM3. This coherency is granular down to the cache line, too. So now every memory allocator and every system call and pointer can be passed directly to the GPU or from GPU to CPU freely.
And the open source driver supports HMM on normal x86_64/aarch64 Linux with consumer-level GPUs today, btw, but it's not as fast or granular. And then there are platforms like Jetson which have used single memory pools for a while; Orin uses a single shared bank of LPDDR5X chips for both CPU and GPU and will get HMM at some point in the future too I assume, though it uses a different driver.
Honestly the only place AMD seems to be winning in terms of compute is on large, bespoke contracts and features like unlocked FP64 performance with parts that are unobtanium and software stacks that have dedicated support engineers. Even Intel seems to be putting up more of a direct fight against Nvidia with oneAPI...
That is the point though, isn't it? Nvidia and AMD are converging to the same model, so it isn't fair to say AMD doesn't understand GPU compute. Nvidia just had a much neater implementation path where they hacked together something that worked in software while their hardware team figured out how to actually implement it. Technically it is arguable that they're behind AMD on general GPU compute, although that'd be pedantic given how thoroughly AMD failed to get their customers a place in the GPGPU market for the last decade.
AMD is floundering, no question. But the failure was understanding the path-dependent implementation aspects. They do understand that GPU compute is essential to the future of computing as an industry. They're clearly putting a lot of resources into that vision and they have been for around 20 years (similar timeline to CUDA).
rocBLAS and other vendor agnostic numeric libraries have made a lot of progress in the past 2 years (mostly as a result of the DoE's exascale computing project)
In fairness, my graphics card isn't supported - multiplying matricies being one of those advanced features that they only implemented in the last couple of years. Older graphics cards maybe don't have the grunt for that.</sarcasm>
I love AMD, the linux graphics drivers are great. But their GPGPU platform is not good.
Now AMDs platform for debugging and profiling GPGPUs apps on the other hand is a different story/mess and very very behind NVIDIAs solutions.
For sure the lack of consumer card support is annoying, all effort seems focused on satisfying their contracts and not expanding support into the much wider GPGPU market rn and I wish it wasn’t. It feels like an afterthought at times. I just wanna be able to compile and play around with HIP on my home computer, but :(
There's some blogs on GPUOpen: MFMA on MI100/200 https://gpuopen.com/learn/amd-lab-notes/amd-lab-notes-matrix...
WMMA on Navi3 https://gpuopen.com/learn/wmma_on_rdna3/
https://www.anandtech.com/show/18721/ces-2023-amd-instinct-m...
> Each A3 supercomputer is packed with 4th generation Intel Xeon Scalable processors backed by 2TB of DDR5-4800 memory. But the real "brains" of the operation come from the eight Nvidia H100 "Hopper" GPUs, which have access to 3.6 TBps of bisectional bandwidth by leveraging NVLink 4.0 and NVSwitch.
Considering that the Xeon Max 9462 is $8000 vs. the H100 going for north of $40,000, that could be interesting.
Socket SP5 has 12 channels, which is 461 GBps per socket at DDR5-4800. Intel is getting 1 TBps from HBM, but then you're paying for HBM. $8000 for the cheapest Xeon Max vs. $3000 for the Epyc 9334 with the same number of cores or ~$1000 for the least expensive thing that will fit in the 12-channel socket. CPUs also have a cost advantage because then you don't need a CPU and a GPU.
Other things might be more compute bound. Then a fast GPU in a socket with a lot of memory channels worth of cheap sticks should be fun.
The physical integrity needed for extremely high bandwidth interfaces is just really tough to achieve on a DIMM-like slot without really advanced high-channel socket topologies. Those numbers listed before aren't for nothing; 2.4TBps bandwith for an 8-socket Xeon vs 2.0Tbps with a 2-socket Xeon using HBM2 is a very significant improvement in overall efficiency.
TPUs were supposed to be their unfair advantage in the cloud ML/DL space. But from what I've experienced, and have heard from other engineers, there's always some subtle incompatibility with TPUs that requires modifying the training/eval scripts. I wonder why they didn't try to polish the rough edges with Pytorch, et al.
If they're admitting TPUs aren't their competitive advantage, then why not sell it to other hosting providers, or hell, even directly to ML scientists and enthusiasts? They'll finally get economies of scale, and take business (and mind share) away from NVidia's monopoly.
And isn't your second paragraph the obvious reason for why this product (A3) exists? It's something they expect to sell to cloud customers who have an existing GPU-based workflow, and just want to run it as-is as fast/cheap/scalable as possible, without worrying about compatibility, and making sure they can always move the workload to some other cloud provider or on-prem if needed.
It's like suggesting Sony releasing some of their games on the PC means they're deprecating Playstation.
(Maybe there would be more details in the IO talk. Does anyone know which one this announcement is from?)
Google owns and designs their own TPUs. They offer these TPUs in the cloud. I've seen many comments in here about how next-level TPUs are (despite zero evidence indicating that). Google even disclaims their TPU by saying that you shouldn't compare it with the H100 given node levels et al.
Their premiere offering is an nvidia H100 offering.
Yes, of course this is a pretty telling indication. If Google was all in on TPUs they'd be building mega TPU systems and pushing those. Instead they're pushing nvidia AI offerings.
A sentence that reads "I am going to eat nothing but vegetables from now on" doesn't mention meat, but you can infer that I won't eat meat again from the sentence.
A sentence that says Google are going all in on nVidea GPUs for AI doesn't need to mention TPUs to convey information about their future either.
Google is huge.
Just a few H100 doesn't represent anything huge in Google scale.
I also tried to find your analogy in that article and google announcement and it's not there.
Where are you reading that Google is going “all in” on nVidia GPUs? I don’t see that in the linked article at all.
These are clearly targeted at their cloud customers who have workloads tailored to GPUs. They’re supplying demand, as cloud providers do.
Companies can do more than thing at a time.
TBF, there is no mention of anything remotely similar to "I am going to eat nothing but vegetables from now on".
If anything, this just reinforces the point I was making. There is nothing at all in the article supporting this narrative. So, where is this coming from? Why are you so intent on this idea that you're reduced to fabricating support for it?
Andnfrom.all that we're meant to say it infers nothing about TPUs?
Offering more options to customers is always better especially when Nvidia has great market share in this area. This is probably the reason why Microsoft is trying to help AMD catch up so their is more competition. AI GPU prices are insane compared to standard GPU because of the lack of competition.
There were articles that Microsoft was helping AMD, but the denied it.
https://arstechnica.com/gadgets/2023/05/microsoft-and-amd-ar...
https://arxiv.org/abs/2304.01433 from April 4 of this year.
> I wonder why they didn't try to polish the rough edges with Pytorch, et al.
It's always funny to me when people have this blindspot - because TPUs aren't for you, they're for the ads org. Neither are PyTorch nor TF for that matter. They're more than happy to get external bug fixers but trust me those individual teams dgaf about external customers. They're not in the least bit community driven projects.
Do you have any more details or links to articles expanding on this?
I don't know of a single GCP product that's been shut down, although I could be missing something. But their track record for GCP is, I think, what you would want a cloud provider's record to be.
(I should mention that I work for GCP. But this is just based on my own memory.)
For example for palm/bard this was my experience:
"Hey we have this amazing LLM!"
"Great, given you are a company can I pay you money above your costs for this service?"
"No but you can register for updates about when the wait-list will open"
They announced cool features for Google docs as well that I can't use.
Some of the things I've seen announced were maybe a year ago and still nothing. Just a wait-list or less.
Not everything gets launched because sometimes they find out in that testing period that they got it wrong.
Side smaller complaint - whats the point in these wait-lists if they never tell me when stuff actually launched.
I haven't seen otter or unicorn models, nor can I find tune them yet.
It comes in 4 sizes, I've only seen two so far
- Immersive mode in Maps (also AI, using NeRF) has only recently added just 5 cities,
- The screenshot-then-Multisearch Near Me is technically shipped, but it seems super-rough; I screenshot my keyboard and it suggested a specific brand of pasta across nearby supermarkets,
- I am still waitlisted for access to LaMDA through the AI test kitchen (and given this year’s I/O, things seem to take a different direction).
There is no question that ChatGPT’s release in particular went by a more successful playbook comparatively.
Just give me a price. Or let me bid on it.
Or, don't announce it like it's launched until it's usable.
Not really relevant then to their announced products is it?
The greatest technological advancement in recent years critically depends on the hardware from a single company with no competition. yet Nvidia stock is still below its 2021 peak. How so?
Perhaps someone like OpenAI has both the expertise and incentive to do so, but not many others.
I'm not saying you are wrong or right, btw
OpenAI is more of a "one-(very impressive)-trick-pony", so they have a stronger incentive.
In a way, it feels like Nvidia is embarrassingly aware of this. They were the reluctant shovel salesman during the cryptocurrency gold rush, and they're rightfully wary of going all-in on AI. If I was an investor, I'd also be quantifying just how much of a "greatest technological advancement" modern machine learning really is.
The cryptomarket was less favorable to Nvidia because it harmed the loyal customers (gamers, AI) for a temporary market (crypto) that indeed largely declined.
As far as stock prices, there was a hype cycle paired with government handouts to the people, these combined to push tech stocks to unreasonable valuations.
Thinking about it, it’s hard to believe how fast the hype cycle moved on from crypto. Only 1-2 years ago every media person, influencer, YouTuber, tweeter etc. were talking about/selling/shilling some kind of crypto, and now all of it seems to have moved on to AGI doomsaying.
Generative AI is used by millions, has very low barrier for entry (it's even free!) and most importantly does not require a network effect so can be valuable immediately.
Surprisingly, they left that out of their sales pitch.
With LLMs everyone+dog is coming out of the woodwork to let people know that it will lead to the extinction of the species.
Not that I don’t think generative AI is a lot more useful than crypto and deserves (some of) the hype. The problem is the hucksters jumping on the hypetrain to continue their $new_hotness grift.
Interesting, the more they warn about it, the more people are eager to invest in it. Kind of a Streisand effect.
Investments should be based on the actual value of the company relative to its price, as well as relative to other investment oppertunities. Trying to making a profit by trading based on historical stock prices will get you whipped by quants who are already doing a much better job of that sort of thing than you could ever hope to do.
It's also not necessarily about the 2021 peak but why isn't Nvidia bigger? allegedly it's a necessary component to a technology that can replace hundreds of millions of people (worth trillions in economic output). And unlike OpenAI, Nvidia wins no matter which company wins the model competition.
https://www.cerebras.net/andromeda/ https://tenstorrent.com/grayskull/
And yet none of them seem to have made any dent in Nvidia’s dominance. None of them have any real presence on industry-standard MLPerf benchmarks (not even TPU releases all benchmarks and they started the damn benchmark).
The truth is that making an AI chip isn’t as simple as putting a bunch of matmuls together in a custom ASIC and pointing a driver at it; there’s hard work and optimization the entire stack down, many of which aren’t even focused on the math part.
So while I don’t doubt that some competitors (AMD?) will gain decent market share eventually, Nvidia’s probably not going to be displaced so easily.
This is why Leading Edge Node will continue to be well funded. Consumer Electronics ( Mainly Smartphone ) Silicon usage has been the main push behind the development of Pure Play leading edge foundry in the past 10 years. Despite the predicted / expected drop of Smartphone sales, considering the potential shown by ChatGPT or Bard, GPU or Wafers dedicated for AI will continue to be in demand for at least another 5 years. In terms of lead time into the investment of silicon development that means we can continue to expect progress all the way till 2030, either 1nm or 0.8nm.
https://arxiv.org/pdf/2303.07470.pdf
https://ieeexplore.ieee.org/abstract/document/9669041
Maybe there will be something like transformers but more suited to crossbar arrays of memristors. If that actually makes sense.
If you’re still doing row-column access, it’s just another Von Neumann machine. If you have compute hardware within each row to perform operations on every row in parallel, it’s now just another ALU.
https://www.hpcwire.com/2023/05/10/googles-new-ai-focused-a3...
https://www.nextplatform.com/2023/05/11/when-push-comes-to-s...
https://www.nextplatform.com/2023/03/21/inside-the-infrastru...
Despite the AI hype, Nvidia’s datacenter revenue was down QoQ and only up 10% YoY.
It remains to be seen if the growth trajectory has changed meaningfully over the last quarter, because the stock is priced for massive earnings growth while their revenue and earnings have been actually shrinking.
We’ll find out on the upcoming earnings call
https://www.macrotrends.net/stocks/charts/NVDA/nvidia/revenu...
https://www.macrotrends.net/stocks/charts/NVDA/nvidia/eps-ea...
Is it? I remember hearing they didn't make much money from their consumer GPU products a few years back. This was one of the reasons why they tried to clamp down so aggressively on people using desktop GPUs for computing. They had made a number of driver changes which restricted the capabilities of anything but the tesla and quadro products. They were also restricting bulk purchases of their cards.
Not that big of a deal these days and I doubt they’d make it so someone couldn’t take an off the shelf GPU and play around with llamas.
Looking right now - I don't see any unbundled Nvidia RTX 4090 cards for sale at Best Buy in New York City that you can go and pick up today. I don't see any desktops with 4090 cards that you can pick up today. I do see one Best Buy in New York City has one laptop with a 4090 card.
Looking at Best Buy in Los Angeles - I see one desktop with a 4090 for sale in West LA that can be picked up today. I don't see any unbundled 4090 cards for sale or laptops with 4090 cards.
I don't know if Nvidia lower end GPUs are sitting on shelves and not selling, but it doesn't look like Nvidia's higher end GPUs are sitting on shelves and not selling.
https://www.techspot.com/images2/news/bigimage/2021/08/2021-...
How could they be sitting on shelves, as they're never put on shelves to begin with, since they're never sold to consumers?
Obviously I was talking about consumer GPUs.
Different workloads require different infrastructure.
Can your workload saturate the TPU without getting throttled by memory or network? Great! Use TPUs and reduce training cost.
But if your TPUs are idle 70% of the time because the constraint is getting data to them ...
"A3 represents the first production-level deployment of its GPU-to-GPU data interface, which allows for sharing data at 200 Gbps while bypassing the host CPU. This interface, which Google calls the Infrastructure Processing Unit (IPU), results in a 10x uplift in available network bandwidth for A3 virtual machines (VM) compared to A2 VMs."
SOTA seems to be 3.2Tbit for H100 clusters so this still seems a bit slow? (Tricky as they don't give us a clear number just 10x). H100's are much more powerful per chip though so at least initially the clusters will be smaller and not network bound.
The tricky thing is no one other than Azure of the big providers seems willing to pay Nvidia's margins for RDMA switches, it seems this is still the case.
The only alternative I could imagine is that TPUs will "win" at supercomputers exclusivity aimed at inference (as opposed to training). Since TPUs excel at inference. The question is how much ML compute is used for inference as opposed to training. Not much, I guess, otherwise something like TPUs would be more popular.
No, it is the 9,163,584th [0] indication that Google likes to pursue multiple solutions in the same space in parallel with different submarkets, risk profiles, expected payoff terms, or other dimensions.
[0] this is a conservative estimate
Clouds offer many competing offerings because different clients have different needs.
Compared to decades ago now everyone carries a supercomputer.
Of course there is always something better on horizon, but if you're building soon that may be worth the wait
Now, the GPU-to-GPU links (NVLink) might often give them a big advantage for some workloads, letting them exchange data without going through the CPU, and virtually address more memory if your want to manipulate very large models.
So it's hard to answer properly without knowing the topology of your cluster.
Also, note that this "supercomputer", is probably "just" a DGX H100 in Google's DC.
First step in doing that is opening up the Android ecosystem and legislating Google's hands out of that pie.
I can't even so much as shit on an android phone without requiring a valid Google account. /crass joke