HPC Systems Special Offer: Two A64FX Nodes in a 2U for $40k
anandtech.com
anandtech.com
... Am I missing something?
Would be good to have some actual benchmarks to compare.
Fugaku cost $1 billion, with 159k nodes. That's only ~$6k per node, and that's if you assume the entire budget is allocated to the compute nodes, which it wouldn't nearly be. Wikipedia lists it as total program cost; so in addition to the facility, cooling and other equipment such as networking, it may well include years worth of operating expenses, staff, etc.
I believe the best way to view the price tag is: low enough that it's no barrier to anyone that's interested in building even a moderate cluster, but high enough to keep away anyone that's just toying around.
That way if somebody buys one of these machines and calls you up to ask a question, you're not wasting your time to treat it as a step in a longer sales process (meaning you let them talk to engineers, respond in depth etc).
but you're probably right that noone buys 2 of these at that price. Probably only japanese companies proud of their country or whatever.
My personal opinion: this is some randomly fixed "we are selling these"-price to show they are not IBMs-"ask for a 6 figure-quote" price. So they might very well be very interesting if you have a budget in the upper six figures (because, the 40k won't be the final offer for sure...); everyone else (like the lower half of the 6-figures) knows: nice, but not my price.
Actually it looks like it's just the opposite of what you describe. Fujitsu build a CPU that's very good at running a wide variety of applications. Unlike x86-64's that need accelerators for good FP or memory bandwidth per node, the A64FX does not need CUDA applications. Plain old fortran is just fine.
Keep in mind the $40k includes support for porting applications, it doesn't mean that if you want 1000 of them that you'll have to pay $20k each.
For another F90 benchmark (Himeno) they got 4 times faster than a dual Xeon 8168 AND 1.1x faster than a Tesla V100 running the CUDA version of the same benchmark. IMO that's pretty amazing, no source code changes and you get BETTER than GPU performance without having to rewrite your code.
So sure a standard Epyc or Skylake refresh will run your code and "just work", the a64fx opens up the possibility of significant performance upgrades with no source code changes. If your code isn't currently rewritten for CUDA getting 10x the bandwidth and 4x the CPU performance would be really attractive.
From wikipedia, a Ford Flex Limited can have a V6 turbocharged 3.6L engine (max torque 350 lb-ft), while a Formula One has a V6 turbocharged with only 1.6L (max torque 214 lb-ft).
The Formula One engine has lower specs, and cost orders of magnitude more.
I think that the answer is that this comparison doesn't make much sense, when the use cases and the target are so different.
Edit: it is funny to see all the replies to my comment mirroring the replies to the parent comment: there is a lot more to the comparison than just cherry-picking a few spec point.
For around the same money you can buy 512 cores of modern AMD x86_64 (4 x 128 core, ~$5000 ea) AND 4 x rtx8000 gpu, 18,000 CUDA compute cores (also ~$5000 ea)
This is likely the development setup for someone that also buys a fully populated rack for production. So it's probably not meant to be a great $$/core story at this small scale.
Though, your point that the AMD EPYC is probably pressuring these ARM HPC setups is fair.
as to the system itself, most of the value in this processor is its memory bandwidth, derived from HBM integration. no other cpu system comes close.
Also note that the #1 on the top 500 list is using the a64fx and scales with 80% efficiency. The #8 on the list is a pure intel (no accelerators) and is 14 times smaller, and only scales with 60% efficiency. That's a pretty impressive feat since good scaling gets harder as the cluster increases in size.
Memory BW? Power consumption ? What's the BW and the power consumption of 2 A64FX CPUs, and what's the BW and power consumption of 4x Threadripper 3990X that also deliver 6 TFLOP/s? You might need to add the power for cooling to the comparison as well if you plan to stock a full cluster with these.
Also, if your main metric is TFLOPs, you might want to add a single A100 to the comparison.... it probably crushes the Threadripper and the A64FX from the POV of TFLOPs, memory BW, and power consumption, but maybe not price.
Your implication is correct however. x86 is more in line with "typical" consumers, and even businesses, who need this kind of compute.
Dell's C6525 quad-node dual-socket EPYC is a good example of what a typical compute-oriented build: https://www.servethehome.com/dell-emc-poweredge-c6525-review...
-------------
A64FX is an HBM2 box. Its a normal CPU (not a GPU), with access to that stupid-high bandwidth RAM. There's probably a few use cases where the high-bandwidth becomes a major advantage.
A64FX compares against the NVidia V100 on a memory-bandwidth and memory-capacity basis (and is approaching GPU-level FLOPs thanks to SVE 512-bit SIMD units). Except its running the ARM instruction set. As others have pointed out, this thing is like Xeon Phi 2.0, except with the notable niche that its the #1 supercomptuer in the world right now.
For example, SVE has mask partitioning and speculative vector load instructions to accelerate data-dependent loop termination. You can do a vector-length-agnostic strncpy on SVE without too much effort.
The basic operation they perform is fast fourier transform on large (number of electrons of the system ^ 2) matrices.
Your 3990X system has dramatically less memory bandwidth with 4 64 bit memory channels good for a peak bandwidth of around 100GB/sec, which is 10% of the A64FX.
Your $3,500 price likely doesn't include an infiniband card, or a port on an IB leaf switch, or the port on spine IB switch, and you are getting quite a bit less memory bandwidth and floating point than the A64FX system. As a result you end up buying more threadripper nodes, paying for more power, more rack space, more cooling, more IB spine switches, more IB cables, more IB leaf switches, and potentially even a larger building ... just to hit the same performance.
Generally to get close to the memory bandwidth or flops with an x86-64 system you end up using an accelerator, like say the Nvidia A100. Problem is that the nvidia GPUs can only run CUDA aware programs, which rules out a significant chunk of HPC workloads.
The attraction of the A64FX is that it's efficient, has impressive flops per watt (which means cheaper cooling, racks, buildings), impressive memory bandwidth, and will run any python, perl, fortran, C, C++, go, java, code you throw at it.
For large clusters this isn't just a nice to have, it's a huge driver of real world performance / price. So much so that the top 4 supercomputers in the world are not x86-64. Even #5 uses a Intel Xeon to run the OS and do I/O and uses an accelerator (Matrix-2000) to do the heavy lifting. #6 and #7 are similar, but using a Nvidia cards. #8 is the first pure x86-64 cluster and is pretty small in comparison, 18 times smaller than #1.
Even #8 (Frontera) being 18 times smaller, makes it MUCH easier to scale codes that run across the entire cluster. Despite that huge advantage, Frontera scaled lipack to the entire cluster with 60% efficiency. The #1 Fugaku using the A64FX scaled at 80% efficiency, that's a pretty large real world advantage.
Keep in mind that Linpack is a pretty limited benchmark, heavily optimized, and not particularly memory or network intensive. Many real world codes will more heavily exercise the system and would likely show a larger performance differential than linpack.
Imagine you have 20,000 watts a rack and 100 racks. Sure you could use cheap nodes, but your total performance would be less... and that's before you throw in problems like scaling to 100 racks at 60% efficiency instead of 80%. When asking for your budget it will be that much harder if you only perform well on CUDA codes. Not that there's not room for cheaper node clusters out there, but non-x86-64 systems do seem to have a significant advantage on larger clusters.
Also general HPC devel/test systems often come with substantial support to enable users to tune their codes to a new platform. I wouldn't assume that if you want 100 racks of them that you'd pay $20k each, in fact the hardware you'd likely use isn't even the same hardware. The A64FX nodes used in clusters like Fugaku have higher density nodes that depend on water cooling, have more cores, and have a much better interconnect.
May be someone could give some hints as to why?
Very interesting memory bus though... 1 TB / s? That is cool, but I would still much rather get a crap load more cores at a reasonable price then be able to send around data that efficiently. Granted, I am definitely not the target audience for this.
HBM not only has lotsa bandwidth(tm) but also much better latency?
My understanding is that, given how memory systems work right now, typically it's the opposite: increasing throughput decreases latency.
in any case, anyone buying one of these is doing so as a poc for a larger procurement.
they're not for tourists.
You would probably buy one of these to facilitate testing and development of codes to run on a big one.
No doubt the memory vs cpu balance is correct for what these computers are doing.
Doesn't matter that none of these critics ever address a single use-case beyond what little info they can gleam from memes. It's particularly obnoxious. But whatever. People who use this tech are all idiots, right, and couldn't possibly know how to add up a balance sheet. Lol, it use lotssa powerr. BRRRRRrrrrrrrr. soo stupid of them to waste that much power for absolutely no reason.
Only partially joking btw!
Especially when you consider that this is a new HPC architecture, that’s not an unreasonable price. As others have mentioned, this is a SKU you’re buying to evaluate a significantly larger purchase.
There isn’t exactly a lot of pricing pressure on the processors. Top end HPC always has the biggest margins because this is a market where absolute speed tends to trump all other factors, including price. A better price comparison might be an IBM Power HPC node — another fairly novel architecture built for raw speed and HPC.
Plus, it is entirely possible that the cost difference of the node (compared to x86) is completely covered by the power savings. HPC is very power hungry. So a significant savings there could make up for an increased node cost.