EU Grabs ARM for First ExaFLOP Supercomputer
hpcwire.com
hpcwire.com
Are the CPUs expected to contribute significant compute, as opposed to marshaling data in/out of the real compute units?
Seems stupid to use millions of dollars of supercomputer time just because you can't be bothered to get a few phd students to spend a few months rewriting in CUDA...
Amortizing the supercomputer over 5 years, a 12 hour job on that supercomputer may cost $63k.
If you want it cheaper, your choices are:
A) run on the supercomputer as-is, and get your answer in 12 hours (+ scheduling time based on priority)
B) run on a cheaper computer for longer-- an already-amortized supercomputer, or non-supercomputing resources (pay calendar time to save cost)
C) try to optimize the code (pay human time and calendar time to save cost) -- how much you benefit depends upon labor cost, performance uplift, and how much calendar time matters.
Not all kinds of problems get much uplift from CUDA, anyways.
I'm curious, what university has a $200MM super computer?
I know governments have numerous Supercomputers that blow past $200MM in build price, but what universities do?
https://www.ncsa.illinois.edu/research/project-highlights/bl...
https://en.wikipedia.org/wiki/Blue_Waters
They have always had a lot of big compute around.
Even when individual universities don't-- governments have supercomputing centers that universities are a primary user of and often charge back value of computing time to the university or it is a separate item that is competitively granted.
Here we're talking about Jupiter, which is a ~$300M supercomputer where research universities will be a primary user.
Basically, unless you have a very specific workload that NVidia has specifically tested, I wouldn't bother with it.
> Seems stupid to use millions of dollars of supercomputer time just because you can't be bothered to get a few phd students to spend a few months rewriting in CUDA...
Rewriting code in CUDA won’t magically make workloads well suited to GPGPU.
Also note — this project is quite modest in scale. Dozens of GenAI clusters larger than this computer will be installed at cloud data centers in the next 18 months.
And
"Jülich is also building out its machine-learning and quantum computing infrastructure, which the supercomputing center hopes to plug in as accelerator modules hosted at its facility."
So a modular setup, where different aspects can be upgraded as needed. Btw:
> Also note — this project is quite modest in scale.
"Exascale" and €273M doesn't sound modest to me. No matter what it's compared against.
If they don’t, then those are big clusters in the sense that AWS is the world’s biggest supercomputer, which is to say, not.
The AI hyperscalers certainly claim to be able to devote 100% of cluster capacity to one training run. Google is training some huge models, OpenAI is also.
The US Dept of Energy had very favorable things to say about the Fujitsu A64FX, which is architecturally similar to the SiPearl Rhea (HBM memory, ARM SVE happy, fast interconnect): https://www.osti.gov/biblio/1965278
They seemed to like the easy porting and flexible programming (since its "just" CPU SIMD) and specifically describe it as competitive with Nvidia:
> To highlight, the pink line represents the energy efficiency metric for A64FX in boost power mode (described in Section IV-C) with an estimated TDP of 140 W and surpassed by the red and yellow lines that represent data for the Volta V100 GPU (highest) and KNL, respectively. The A64FX architecture scores better with the energy efficiency metric relative to the performance efficiency metric due to its low power consumption.
In fact, ARM A64FX supercomputers topped the Green500 for some time, which is the global supercomputer power efficiency ranking, outclassing Nvidia/Intel/AMD machines.
High perf/watt matter more than just high perf/node, but even that balanced against 'how low latency can the interconnect be'.
You then hit the high FLOP count with tons of nodes.
To be fair Nvidia realized this paradigm years ago too, which is why they bought Mellanox.
So a lot of factors come out of that, and a lot of designs that take interesting stabs at new balances towards that goal.
typically at least in the US there's a mix of GPU-focused machines as well as traditional CPU-focused machines. the leadership class machines (i.e., the machines funded to push the FLOPS records) tend to be highly focused on GPU. one reason is fixed cooling/power availability. I assume these facilities are looking at ARM as a way to save 10-20% on power and thus cram that much more into the facility.
https://www.eenewseurope.com/en/sipearl-raises-e90m-for-rhea...
In any case, SiPearl seems to be the one designing the actual chip, they are French.
In general, I think ARM just was famously started in the UK and so they’ll always be associated with the country in some intangible way.
It is sort of funny that we label companies like this, really they are all multi-national entities. Especially in the case of a company like ARM—they license out the designs to be (sometimes quite significantly!) customized by engineers in other countries, and I’m sure they integrate lots of feedback from those partners. Then those designs are often actually fabricated in a third country!
Which is good, the world is best when we all need each other.
What a weird metric for determining the nationality of a company. Intel are publicly traded: are they stateless?
The chip is being designed by a French company, they can license the IP from outside the EU while still building up the EU domestic chip building capabilities. They’ve just outsourced one (big) piece of the puzzle.
Calling ARM a Japanese company was just to highlight the international nature of these sorts of projects.
Interesting take on geography .
The confusion, pobably has to do with the fact that the German tier 0 Gauss super computing center is actually spread over 3 sites (Jülich near Cologne/Aachen, Stuttgart and Garching near Munich)
This reads weird. It took me way too many seconds of wondering "wouldn't Stuttgart be nearer to... Stuttgart?" before I understood what you wrote. Sometimes the Oxford Comma has value, it seems.
You don't want something like that in a city centre.
Nuclear sites also have the tendency to be built on a nation's border.
https://www.anandtech.com/show/16072/sipearl-lets-rhea-desig...
https://semiengineering.com/tag/sipearl/
...It may even be behind schedule?
https://fuse.wikichip.org/news/3256/centaur-new-x86-server-p...
https://fuse.wikichip.org/news/3099/centaur-unveils-its-new-...
Imagine if it came out today. I feel like its the near perfect architecture for cheap GenAI.
Looked neat on paper, but paper is just that at the end...
> SiPearl chose ARM as it is well-established and ready for high-performance applications. Experts say RISC-V is many years away from mainstream server adoption.
China would benefit much more than US/Europe from RISC-v catching up. Wouldn’t be the smartest thing to do geopolitically (or longterm economically for that matter).
An ARM and a leg, for sure.
HPC contracts are generally borne of federal-agency RFPs, and are extremely competitive, and they only 'pay out' upon a passed acceptance test, so it's not trivially possible to predict which quarter your revenue will land for a given sale. You wind up with sales teams putting tons of work into a contract that didn't get selected, which sucks, but even if you win you might wind up missing sales goals, and then overshooting the mark the following quarter.
In a company less hidebound this obviously wouldn't be a problem, but IBM has been run by the beancounters for long enough that the prestige isn't worth the murky forecast.
The article focuses basically on the x86 vs. arm competition.
Any idea where to read more about the application this machine is expected to run? I guess the usual like weather forecast and such?
Does NVIDIA even sell those anymore without the whole package deal since they came up with Grace?
The last supercomputer with NVIDIA GPUs and third party CPUs I remember reading about was with Zen 2 cores, multiple years ago.
But yes, the CPU is mostly just a footnote, most of the FLOPs come from the GPUs. Although of course the CPUs still need to be sufficiently fast enough that the GPUs can be kept fed.
IIRC, on Perlmutter's GPU partition, 60 of its 64PFLOPs are represented by the GPUs, with the remaing 4PFLOPs, being the CPUs. In comparison, their previous system Cori, had ~3PFLOPs in the Haswell partition and ~30PFLOPs on the KNL partition.
That, to me, indicates that CPU performance is mostly a footnote when GPUs are involved, as anyone who had previously been using Cori and is now on Perlmutter, will not see as dramatic of an improvement if restricted to CPUs but they would if able to use GPUs.
From where I sit, we're often limited by memory bandwidth. When a CPU such as A64FX or even M2 shows up with decent bandwidth, lo and behold, they are often competitive. I do not understand why we didn't see something like SPR Max years ago.
And regarding your question about GPU/accelerators, CPUs still do a LOT of work in HPC. I'm guessing they chose ARM for performance per watt, very important when scaling to many processors.
What would a healthy EU HPC ecosystem look like? At some point there was some excitement about Beowulf clusters [1]. When building a new supercomputer, think, for example about making at least its main compute units more widely available (universities, startups, SME's etc). HPC Computing is arcane and to tap its potential in the post-Moore's law era it needs to get much more democratized and popular.
If they did, would anybody want them? Are those units competitive for smaller setups and the kind of jobs they run?
* at least if they used pci-e
on the second branch of your question, indeed a local "supercomputer piece" should have a sufficient number of CPU/GPU's to pack meaningful computational power. this way it would also require and enable the right kind of tooling and programming that scales to larger sizes.
given that algorithms can enhance practically any existing application (productivity, games etc), this might be a case of "build it and they will come"
[1] https://www.pcmag.com/news/intel-ceo-get-ready-for-the-ai-pc
"Since 2017, every system on the Top500 list of the world's fastest supercomputers has used Beowulf software methods and a Linux operating system."
As for accessible by everyone: here is how you can apply for computing time via PRACE, if you work at an academic institution, a commercial company or a government entity located in Europe:
https://prace-ri.eu/call/eurohpc-ju-call-for-proposals-for-r...
In addition to the very large machines that are covered by PRACE, typically there are national calls for access to "smaller" HPC resources, say up to a few million CPU-hours per year. The allocations on PRACE average around 30-40 million cpu-hours.
What is explicitly NOT allowed on these machines is typically running jobs that use just a handful of cores. They've paid a lot of money for the fancy interconnect, amd they want to see it used.
It already exists, there are probably >100 HPC clusters spread throughout the EU in universities + the CERN cluster etc.
> startups, SME's etc
Why would we want to provide resources for a startup to waste compute resources to optimise advertising clicks? They can spend their VC cash at aws.
ultimately this is also better use of taxpayer money: diffusing technology more wider and educating people to make use of supercomputing technologies beyond the ivory towers
But if we are gone do a HPC thing, at least make the processor open-hardware and RISC-V.
Doubly so on the consumer side of things.
RISC-V is coming, it just takes a long time.
On one hand, it's nice to see funding to european companies to develop european technolgy, aiming at a technological sovereignty.
On the other hand, SiPearl looks like it was virtually unknown up to this point, and I can't seem to find anything looking like a cpu review (their website claims they have already released at least one generation of Rhea cpus). So this amount of money might not be wasted but still not optimally spent. Which isn't 100% bad, but at least bittersweet.
If anything, without reviews and performance benchmarks, we might just get ExaFLOPS on paper.
Like how does one of these Rhea CPUs compare to, say, a Graviton 2/3 or to an Ampere Altra cpu?
https://www.anandtech.com/show/16072/sipearl-lets-rhea-desig...
(Neoverse V1 and HBM2e woild make this chip kinda old when its finally operational).
CPU design takes many years, and this was a HPC only chip, so it doesn't necessarily need to be marketed and paraded around, and the workloads will be totally different than what Graviton processors run.
Like many vendors in the ARM space, most of the real innovation and design comes from ARM.
I'll chuck out an unqualified estimate of €10k each, will find out next year (probably) if I'm anywhere close!