Are the CPUs expected to contribute significant compute, as opposed to marshaling data in/out of the real compute units?
Are the CPUs expected to contribute significant compute, as opposed to marshaling data in/out of the real compute units?
The US Dept of Energy had very favorable things to say about the Fujitsu A64FX, which is architecturally similar to the SiPearl Rhea (HBM memory, ARM SVE happy, fast interconnect): https://www.osti.gov/biblio/1965278
They seemed to like the easy porting and flexible programming (since its "just" CPU SIMD) and specifically describe it as competitive with Nvidia:
> To highlight, the pink line represents the energy efficiency metric for A64FX in boost power mode (described in Section IV-C) with an estimated TDP of 140 W and surpassed by the red and yellow lines that represent data for the Volta V100 GPU (highest) and KNL, respectively. The A64FX architecture scores better with the energy efficiency metric relative to the performance efficiency metric due to its low power consumption.
In fact, ARM A64FX supercomputers topped the Green500 for some time, which is the global supercomputer power efficiency ranking, outclassing Nvidia/Intel/AMD machines.
Seems stupid to use millions of dollars of supercomputer time just because you can't be bothered to get a few phd students to spend a few months rewriting in CUDA...
Amortizing the supercomputer over 5 years, a 12 hour job on that supercomputer may cost $63k.
If you want it cheaper, your choices are:
A) run on the supercomputer as-is, and get your answer in 12 hours (+ scheduling time based on priority)
B) run on a cheaper computer for longer-- an already-amortized supercomputer, or non-supercomputing resources (pay calendar time to save cost)
C) try to optimize the code (pay human time and calendar time to save cost) -- how much you benefit depends upon labor cost, performance uplift, and how much calendar time matters.
Not all kinds of problems get much uplift from CUDA, anyways.
I'm curious, what university has a $200MM super computer?
I know governments have numerous Supercomputers that blow past $200MM in build price, but what universities do?
https://www.ncsa.illinois.edu/research/project-highlights/bl...
https://en.wikipedia.org/wiki/Blue_Waters
They have always had a lot of big compute around.
Even when individual universities don't-- governments have supercomputing centers that universities are a primary user of and often charge back value of computing time to the university or it is a separate item that is competitively granted.
Here we're talking about Jupiter, which is a ~$300M supercomputer where research universities will be a primary user.
Basically, unless you have a very specific workload that NVidia has specifically tested, I wouldn't bother with it.
> Seems stupid to use millions of dollars of supercomputer time just because you can't be bothered to get a few phd students to spend a few months rewriting in CUDA...
Rewriting code in CUDA won’t magically make workloads well suited to GPGPU.
High perf/watt matter more than just high perf/node, but even that balanced against 'how low latency can the interconnect be'.
You then hit the high FLOP count with tons of nodes.
To be fair Nvidia realized this paradigm years ago too, which is why they bought Mellanox.
So a lot of factors come out of that, and a lot of designs that take interesting stabs at new balances towards that goal.
typically at least in the US there's a mix of GPU-focused machines as well as traditional CPU-focused machines. the leadership class machines (i.e., the machines funded to push the FLOPS records) tend to be highly focused on GPU. one reason is fixed cooling/power availability. I assume these facilities are looking at ARM as a way to save 10-20% on power and thus cram that much more into the facility.
Also note — this project is quite modest in scale. Dozens of GenAI clusters larger than this computer will be installed at cloud data centers in the next 18 months.
And
"Jülich is also building out its machine-learning and quantum computing infrastructure, which the supercomputing center hopes to plug in as accelerator modules hosted at its facility."
So a modular setup, where different aspects can be upgraded as needed. Btw:
> Also note — this project is quite modest in scale.
"Exascale" and €273M doesn't sound modest to me. No matter what it's compared against.
If they don’t, then those are big clusters in the sense that AWS is the world’s biggest supercomputer, which is to say, not.
The AI hyperscalers certainly claim to be able to devote 100% of cluster capacity to one training run. Google is training some huge models, OpenAI is also.