Intel Gaudi 3 the New 128GB HBM2e AI Chip in the Wild
servethehome.com
servethehome.com
An interesting aspect of Intel's design is they use Ethernet for connectivity. If they can get the performance on par with NVLink, that by itself could be a win because everybody knows how to manage Ethernet. Very few people know how to manage an NVLink network.
To be clear, this is data center hardware. The lower power versions of these cards consume like 600W, and no mention in the article on pricing.
Ethernet has so, so many gotchas. Maybe if it was a layer 3 only network it would work. Maybe.
Edit: and yes, it's RDMA over ethernet https://docs.habana.ai/en/latest/Gaudi_Overview/Gaudi_Archit...
NVLink either talks native NVlink to itself when you are using NVlink switches either intra-server or intra-rack or;
It can talk PCIe over NVlink when talking to a PCIe endpoint.
Or you can run Infiniband or Ethernet on top of it and talk to w/e is on the other side.
Gaudi isn’t that different remember Ethernet != TCP/IP.
In my understanding one of the big advantages of the protocol (v2, that is) is that it is routed over IP and can work with existing switches ($$) instead of needing specialized ones ($$$$)
When these Intel GPUs are “in the wild” it actually means Xeon salespeople are out on the hunt.
Interestingly, tenstorrent is doing something similar with their wormhole cards.
I'm not convinced yet that it is the right way to go. If the switching fabric on the card fails, you lose the whole card. Keeping it separated out is a bit less risky, at the cost of some speed.
I'm more partial to composable fabrics, but they aren't ready yet for PCIe5 and we have PCIe6 just around the corner next year.
What I'd prefer is the connection is through the UBB/OAM baseboard, such that you have PCIe connections. Look into what GigaIO and Liqid are doing. There is a 3rd option that is even cooler than those two, but I don't want to mention it here. ;-)
But it’s a serious thing if it happens.
The CEO of NVIDIA was pretty clear: Nvidia's focus is now on selling Blackwell AI Factories. According to some napkin math, each of these Blackwell AI factories will have an annual power bill of roughly $8,000,000 USD and if you do a little apples to oranges comparison will outperform the computer that's currently #1 on the supercomputing top 500 (Frontier at Oak Ridge) by one to two orders of magnitude.
for comparison, over the past fifteen years, we've seen 3 orders of magnitude improvement in the supercomputing top 500. (from roadrunner in 2008 to frontier in 2024). Nvidia is going to do 1 to 2 orders of magnitude improvement instantaneously.
Nvidia's core market has shifted dramatically. There is no competition.
In the keynote Nvidia is quoting FP4 performance for Blackwell specifically. Frontier and every computer that has ever hit the supercomputer top 500 is measuring fp64. This is why I qualified my statement above by saying if you’re willing to compare apples and oranges.
But I think this is where things open up to debate (and/or personal interpretation). My view is, if all you care about is AI workload (Specifically LLMs), then you really are seeing a full two orders of magnitude improvement. And if we are trying to get a feel for the space, and what that even means, then there really is nothing else to compare Blackwell to other than frontier, in terms of scale alone.
Above I say “one or two orders of magnitude”, the “one order of magnitude” is a value I got by taking Blackwell’s FP4 values and dividing by 16, to try to conceptually convert back to a value that can be compared to fp64. — I’m perfectly happy to admit that all I’m really capable of doing here is using estimates to compare a completely new thing (the new compute architecture of Blackwell) to what came before it (the majority of the history of the supercomputing 500). But that’s okay, because that’s my only goal.
—-
Believe me, I get it. Blackwell will never post a linpack score. Blackwell will never run big scientific computing jobs like… global weather simulations.
But as computer science nerds, how could we not be excited to see such a big move: changing the computer architecture to better suit a specific compute workload (LLMs)?
I know Bitcoin has its asic processors, but to me Blackwell’s mission statement of “computing intelligence” is a lot more interesting.
By my count, 49 corporate logos were shown as launch partners.
At one point in the talk the CEO claims that “Blackwell will be Nvidia’s most successful product ever”.
—
And again, the product he’s talking about is a data centers packed with 56 racks…
Is it inevitable? I think so. Before 2019 there wasn't an opportunity, now there is.
For software, Chinese universities, Alibaba, Tencent and Bytedace are already releasing models, training code and in rare cases datasets that are competitive with private offerings. CogVLM/CogAgent is one that I use. It's very promising.
But, anyway, we will prohibited from buying it, probably. We still can't buy Cuban cigars.
C'mon, do something...
Wow and I thought that the latest generation of GPUs was better.
With the state of OpenCL it's frustrating. So much to be had, but so little improvement and support.
Frustrating is a good way to put it, but there's a certain causal satisfaction I feel from watching it all unfold. Of course everyone loses to CUDA when they refuse to sponsor an Open Source alternative. It's fascinating to me that hardware manufacturers would rather let CUDA dominate than establish a basic working relationship.
Additionally, due to the KYC requirements around these GPUs (due to US export controls), we really want to get to know our customers first.
Feel free to ping me on email and happy to get on a call and talk more.
The problem I realized over a year ago was that nobody had hourly rental access to high end AMD GPUs. In addition, access to high end Nvidia was equally difficult. I signed up for a CoreWeave account, put in my credit card and was told a few weeks later that my account was not approved.
In effect, the only way to get access to super high end compute, was to be involved in HPC and that requires connections. At the time, we also didn't even know if AMD was going to seriously adopt AI as a strategy.
My view was that there were actually two problems, lack of general access and that everyone was putting all their eggs into a single basket. Mostly because of that lack of access, and because AMD was lacking a great developer flywheel story.
I spent August to December building a business plan, closing funding, forming the business, hiring my co-founder full time, securing data center space, securing direct relationships with vendors, and designing the system we were going to deploy. There are a million other little details in there, but this is long enough as it is.
Oct/Nov of last year rolls around and suddenly AMD has changed their tune. Lisa Su doubles down. Dec 6th, MI300x rolls out. We made our first PoC order in January, received it in March. It just goes to show how cutting edge and how long all of this takes. 3 more small (not hyperscaler) businesses sprung up during that time, all offering effectively the same product. We went to the data center, deployed our PoC and about 2 weeks later, we had our first customer onboarded. I call all of that validation, and was able to secure further funding based on it.
To answer your question, I'm not sure that I need a specific USP. The demand for compute isn't going down. If I have a product that people want, and I can offer them ethical, honest, truthful, great service around that product. All based on decades of experience. Can't that be enough? Myself and my investors believe so.